Yunqing Zhao

Senior AI Researcher at Tencent Hy. I received my PhD (Outstanding Thesis Award) from SUTD, 2024, advised by Prof. Ngai-Man Cheung. I was a Senior Research Scientist at TikTok / ByteDance, and before that a research intern at Sea AI Lab, Microsoft Research Asia, and ByteDance AI Lab.

One thread runs through my work: context. I study how models learn from very little of it 01, how it can be turned against them 02, how they learn to build their own 03, and how RL teaches agents which context to seek 04. Each gets its own chapter below.

My work has been published at NeurIPS, ICML, ICLR, CVPR, ICCV, and ECCV, and in TIP and TPAMI.

Yunqing Zhao

Research

Weights hold what a model knows. Context decides what it does.

A trained model is a fixed prior. Everything it does afterwards (adapting to a new domain, following a prompt, answering a question about an image, calling a tool) is conditioned on a context. In-context learning made this explicit: same weights, different context, different behavior. Many failures are failures of context. With too little, a model overfits. With the wrong kind, it can be steered. And the right kind is often not given at all; it has to be gathered.

My research has tracked the context as it moved: from a few examples that had to be written into the weights, to an open input that anyone can tamper with, to evidence the model assembles for itself, and now to context an agent learns to seek.

The chapters below trace that path.

Sought context 2026 →

Multimodal LLMs, agents & media generation

Teach it which context to seek. Agents that go and get what they need.

CodeDance used RL to teach a multimodal model when a tool is worth calling; ThinkGen used it to teach one model to write another’s instructions. At Tencent Hy I work on multimodal LLMs and agents with a data-centric approach: teaching them what to look at, which tool to call, and when the evidence they have gathered is enough and accurate.

Some of this work shows up in HyCreator, a long-video agent that turns a short brief into a film. Before it assembles picture and sound, it creates references for each character, location and prop, asks a subagent for an independent review, and checks every clip for continuity.

Now · Senior AI Researcher, Tencent Hy · hy.tencent.ai
Constructed context 2025–26

Multimodal LLMs & tool use

Let the model build its context. Pixels, time and code, turned into evidence it can read.

Today’s multimodal models don’t just read their context; they have to build it. Pixels must become tokens a language model can read and write: a vision encoder pretrained to speak (GenLIP), and segmentation masks written as compact text (Text4Seg++). A long video mixes timestamps, saliency and captions, so each kind of token gets its own expert (TimeExpert). One model can even compose another’s context: an MLLM reasons first, then writes the instruction a diffusion transformer draws from (ThinkGen). And when one look isn’t enough, the model writes and runs code that draws boxes, lines and plots back into its own context, with a reward that keeps it from calling tools it doesn’t need (CodeDance). This body of work also contributed training data and evaluation to Pistis, a family of 27B and 9B multimodal models.

Evidence on demand CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning Executable code as a general visual solver: it composes tools and renders evidence it can check. A balanced tool-calling reward curbs overuse, and unseen tool compositions emerge during RL. CVPR 2026 · tech lead, corresponding author Vision that speaks Let ViT Speak: Generative Language-Image Pre-training Pretrain the ViT to emit language tokens directly, so its features arrive in the language model’s own terms. ECCV 2026 · co-author Masks as text Text4Seg++: Advancing Image Segmentation via Generative Language Modeling Extends Text4Seg (ICLR 2025): box-wise semantic descriptors turn masks into “semantic bricks”, so segmentation becomes next-brick prediction, with no mask decoder. TPAMI 2026 · tech lead, corresponding author Routing time TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding A mixture-of-experts Video-LLM that routes each task’s tokens to a specialist, reaching state of the art on dense captioning, moment retrieval and highlight detection. ICCV 2025 · co-author Think, then draw ThinkGen: Generalized Thinking for Visual Generation Chain-of-thought for image generation: the MLLM plans, the diffusion transformer draws, and alternating RL (SepGRPO) trains both. CVPR 2026 · co-author Data & evaluation Pistis Technical Report 27B and 9B multimodal models trained with IDRL — alternating on-policy distillation and RL in one loop — for stronger reasoning and long-horizon agentic tasks. Technical Report · training data & evaluation
Untrusted context 2023-2024

AI safety for multimodal models

Guard what the model reads. An open input is an attack surface, and a key.

If a context can teach a model, it can also mislead it. Once large models read images and prompts at inference time, the context became an interface, and whoever writes into it can steer the model. With only black-box access, an adversarial image makes open VLMs such as MiniGPT-4 and LLaVA give the attacker’s chosen response (AttackVLM). The same channel can serve the owner: a trigger prompt only they know makes a watermarked Stable Diffusion draw a predefined watermark, such as a scannable QR code, while other prompts behave as before (WatermarkDM, US patent).

Scarce context 2022–23

Few-shot learning & knowledge transfer

Learn from a little context. Fewer examples, so the prior matters more.

In few-shot image generation, a new domain arrives as about ten images, and the only place to put them is the weights. Adaptation becomes a negotiation between context and prior: what should a pretrained model keep, and what should the few examples overwrite? We found that it mostly fails by collapsing onto the examples (A Closer Look), that the prior worth keeping depends on the target (AdAM), and that some of it has to be actively forgotten (RICK). FS-BAN asks the same question of classifiers facing domains they have never seen. In-context learning answers it without gradients; this was the gradient version.

Prelude · 2020

Two seeds. Explanation-guided training helped cross-domain few-shot classifiers generalize from a handful of examples (ICPR, code) 150+ citations, and a sparse adversarial attack on object detectors placed 10th of 1,701 in the CIKM AnalytiCup (paper, code). The first grew into Chapter 01, the second into Chapter 02.

Experience

Where the chapters were written.

Senior AI Researcher, Foundation Model Dept.
Jun 2026 – Present
TikTok / ByteDance, Singapore03 · Constructed context
Senior Research Scientist Jan – Jun 2026Research Scientist Jul 2024 – Dec 2025
With Song Bai and Zilong Huang. Contributed training data and evaluation to Pistis (Technical Report), a family of 27B and 9B multimodal models.
Jul 2024 – Jun 2026
Microsoft Research Asia
Research Intern.
2023–2024
Sea AI Lab, Singapore02 · Untrusted context
Research Intern
2022–2023
ByteDance AI Lab, Singapore01 · Scarce context
Research Scientist Intern
2021–2022
ST Electronics-SUTD Cyber Security LabPrelude
PhD Student Researcher
2020–2021
University of Hong Kong
MS Student & Research Assistant
2018–2019

Publications

Selected papers, one thread.

† corresponding  ·  * equal contribution  ·  Full list on Google Scholar

CodeDance
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
Qi Song*, Honglin Li*, Yingchen Yu, Haoyi Zhou, Lin Yang, Song Bai, Qi She, Zilong Huang†, Yunqing Zhao†
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Summary
Do we always need to integrate tools for visual reasoning? While tool-augmented MLLMs have shown strong gains, indiscriminate tool invocation results in unnecessary computation and reasoning inefficiency. CodeDance learns when and how to use tools – and critically, when tool invocation is unnecessary. Through a difficulty-aware RL objective (RBAT), it achieves strong improvements across counting, chart QA, and visual search/math benchmarks while reducing reasoning turns.
Text4Seg++
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
Mengcheng Lan, Chaofeng Chen, Jiaxing Xu, Zongrui Li, Yiping Ke, Xudong Jiang, Yingchen Yu, Yunqing Zhao†, Song Bai
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026
Summary
Can a multimodal LLM segment at the pixel level without bolting on a separate mask decoder? Text4Seg++ recasts segmentation as pure text generation via semantic descriptors that map each image patch to a text label, with a Row-wise Run-Length Encoding (R-RLE) that trims sequence length by 74% and speeds inference 3x. Reframing the task as next-brick prediction, it outperforms state-of-the-art models across natural and remote-sensing benchmarks with no task-specific fine-tuning, while staying compatible with existing MLLM backbones.
Pistis
Pistis Technical Report
Heyun Chen, Xiaohan Lan, Jiaxi Li, …, Xudong Zhang, Yunqing Zhao, Shuai Zheng
Technical Report, 2026
Summary
Pistis is a family of 27B and 9B multimodal models built on Qwen3. Post-training uses Interleaved Distillation and RL (IDRL) – alternating on-policy distillation and RL within a single loop rather than sequentially – for stronger reasoning and long-horizon agentic tasks. Contribution: training data construction and evaluation.
GenLIP
Let ViT Speak: Generative Language-Image Pre-training
Yan Fang*, Mengcheng Lan*, Zilong Huang†, Weixian Lei, Yunqing Zhao, Yujie Zhong, Yingchen Yu, Qi She, Yao Zhao, Yunchao Wei†
European Conference on Computer Vision (ECCV), 2026
Summary
What if a vision encoder were pretrained the way LLMs are – by simply predicting the next token? GenLIP makes the ViT itself generative: it reads image patches and emits language tokens directly under a plain language-modeling objective, dropping the contrastive two-tower setup and any separate text decoder. The result is a strikingly simple recipe that aligns visual features with the autoregressive nature of LLMs and transfers cleanly into multimodal models.
ThinkGen
ThinkGen: Generalized Thinking for Visual Generation
Siyu Jiao*, Yiheng Lin*, Yujie Zhong†, Qi She, Wei Zhou, Xiaohan Lan, Zilong Huang, Fei Yu, Yingchen Yu, Yunqing Zhao, Yao Zhao, Yunchao Wei†
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Summary
Chain-of-thought reasoning has reshaped multimodal understanding, but in image generation it has stayed tied to specific scenarios. ThinkGen pairs a pretrained MLLM, which reasons about the user’s intent and writes a tailored instruction, with a Diffusion Transformer that generates from it. A separable GRPO paradigm (SepGRPO) alternates reinforcement learning between the two modules, so they can be trained jointly across diverse datasets; the model reaches state-of-the-art results on multiple generation benchmarks.
TimeExpert
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
IEEE/CVF International Conference on Computer Vision (ICCV), 2025
Summary
Video temporal grounding bundles three distinct subtasks – localizing timestamps, scoring saliency, and generating text – yet most Video-LLMs push them all through identical pathways. TimeExpert is a Mixture-of-Experts Video-LLM that routes each task-specific token to a specialized expert, sharpening event modeling and efficiency to reach state-of-the-art on dense video captioning, moment retrieval, and highlight detection.
WatermarkDM
A Recipe for Watermarking Diffusion Models
Technical Report, 2023  ·  US Patent, 2024 250+ citations
Summary
We derive a recipe for efficiently watermarking state-of-the-art diffusion models (e.g., Stable Diffusion), via training from scratch or fine-tuning. Our recipe is straightforward but involves empirically ablated implementation details, providing a solid foundation for future research on watermarking DMs.
AttackVLM
On Evaluating Adversarial Robustness of Large Vision-Language Models
Conference on Neural Information Processing Systems (NeurIPS), 2023 600+ citations
Summary
Large VLMs achieve unprecedented performance with visual inputs, but multimodal generation exacerbates safety concerns. We evaluate the adversarial robustness of open-source VLMs (e.g., MiniGPT-4, LLaVA, BLIP, UniDiffuser) in the most realistic high-risk setting, where adversaries have only black-box access and seek to elicit targeted responses.
RICK
Exploring Incompatible Knowledge Transfer in Few-Shot Image Generation
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
Summary
Through interpretable GAN dissection, we show that fine-tuning-based methods cannot effectively remove knowledge incompatible with the target domain after adaptation. We propose RICK, an efficient algorithm that estimates filter importance and prunes those incompatible with the target domain for few-shot image generation.
FS-BAN
FS-BAN: Born-Again Networks for Domain Generalization Few-shot Classification
Yunqing Zhao, Ngai-Man Cheung†
IEEE Transactions on Image Processing (TIP), 2023
Summary
We propose a method to improve generalizability for cross-domain few-shot classification using born-again networks. Our algorithm requires no additional parameters or training data and can be readily applied to existing FSC models, distilling dark knowledge from a teacher via multi-task objectives designed for cross-domain few-shot learning.
AdAM
Few-Shot Image Generation via Adaptation-Aware Kernel Modulation
Conference on Neural Information Processing Systems (NeurIPS), 2022
Summary
When fine-tuning a pretrained generator on few-shot target samples, we show that state-of-the-art algorithms perform no better than a simple baseline when the target is distant from the source domain. We propose AdAM, a parameter-efficient and target-aware method to select source knowledge important for few-shot adaptation.
LS-KD
Revisiting Label Smoothing & Knowledge Distillation Compatibility: What was Missing?
International Conference on Machine Learning (ICML), 2022 Spotlight
Summary
Prior work disagreed on whether a label-smoothed teacher helps or hurts knowledge distillation. We pinpoint systematic diffusion as the missing piece – it erodes the benefits of distilling from an LS-trained teacher at high temperatures – and show that pairing an LS-trained teacher with low-temperature transfer reliably produces strong students.
FSIG
A Closer Look at Few-shot Image Generation
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022 100+ citations
Summary
We analyze existing few-shot image generation algorithms in a unified testbed and find that diversity degradation is the major issue during few-shot target adaptation. Our mutual information based algorithm alleviates this issue and achieves state-of-the-art performance on few-shot image generation tasks.

Service

Teaching & reviewing.

Reviewer NeurIPS, CVPR, TPAMI, TIP, TIFS, TNNLS, TASL, TMM, TCSVT, CVIU.
Teaching Assistant 50.021 Artificial Intelligence and 50.035 Computer Vision at SUTD.