Multimodal LLMs, agents & media generation
Teach it which context to seek. Agents that go and get what they need.
CodeDance used RL to teach a multimodal model when a tool is worth calling; ThinkGen used it to teach one model to write another’s instructions. At Tencent Hy I work on multimodal LLMs and agents with a data-centric approach: teaching them what to look at, which tool to call, and when the evidence they have gathered is enough and accurate.
Some of this work shows up in HyCreator, a long-video agent that turns a short brief into a film. Before it assembles picture and sound, it creates references for each character, location and prop, asks a subagent for an independent review, and checks every clip for continuity.




