I am a Ph.D. student in Electrical and Computer Engineering at The University of Hong Kong, advised by Prof. Xihui Liu at HKU MMLab. I received my B.Eng. from Zhejiang University, where I worked with Prof. Chunhua Shen and Prof. Hao Chen at the State Key Laboratory of CAD&CG.
My research interests lie in generative models, unified multimodal models, and agents. I am particularly interested in scalable visual generation, long-context multimodal learning, and generalist agents that can perceive, reason, and act in real-world environments.
I am currently a research intern in the Tencent Hunyuan Qingyun Program. Previously, I worked at Meituan with Manyuan Zhang as my mentor, and at Alibaba Tongyi Wan, where Ruihang Chu mentored me on video generation research.
🔥 News
- Jul. 2026. 🤖 We release UniClawBench, a universal benchmark with 400 bilingual real-world tasks for proactive agents.
- Jul. 2026. 💼 I join the Tencent Hunyuan Qingyun Program as a research intern.
- Mar. 2026. 🎨 We release MACRO, featuring MacroData, a 400K structured long-context dataset for multi-reference image generation.
- Dec. 2025. 🎬 We release Wan-Move, a scalable framework for precise motion-controllable video generation.
- Sep. 2025. 🎓 I start my Ph.D. in ECE at The University of Hong Kong and receive the HKU Presidential PhD Scholar Programme award.
- 2025. 🎉 TTS-VAR and Wan-Move are published at NeurIPS 2025.
- 2025. 🎉 Framer is published at ICLR 2025.
- Jun. 2025. 🎓 I receive my B.Eng. from Zhejiang University.
- 2024. 🎉 FreeCompose is published at ECCV 2024.
📈 Citations
Updated from Google Scholar · July 23, 2026
Citation data is temporarily unavailable.
📝 Selected Publications
Selected publications are listed in reverse chronological order by first public submission. Please see Google Scholar for the complete list.

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks \ Preprint, 2026 \ Zhekai Chen*, Chengqi Duan*, Kaiyue Sun*, Bohao Li, Yuqing Wang, Manyuan Zhang†, Xihui Liu†
A capability-driven benchmark containing 400 bilingual, real-world tasks for proactive agents. Evaluates skill usage, exploration, long-context reasoning, multimodal understanding, and cross-platform coordination.

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data \ Preprint, 2026 \ Zhekai Chen, Yuqing Wang, Manyuan Zhang†, Xihui Liu†
Introduces MacroData, 400K structured samples with as many as ten reference images. Covers customization, illustration, spatial reasoning, and temporal prediction under long multimodal context.

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance \ NeurIPS 2025 \ Ruihang Chu*†‡, Yefei He*, Zhekai Chen*, Shiwei Zhang†, Xiaogang Xu, Bin Xia, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu, Yingya Zhang, Yujiu Yang†
Enables precise object and camera motion control using latent trajectory guidance. Integrates motion-aware conditions into an off-the-shelf image-to-video model without architectural changes.

TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation \ NeurIPS 2025 \ Zhekai Chen, Ruihang Chu†, Yukang Chen, Shiwei Zhang, Yujie Wei, Yingya Zhang, Xihui Liu†
The first general test-time scaling framework for visual autoregressive models. Formulates generation as path search with diversity exploration and potential-based selection.

Framer: Interactive Frame Interpolation \ ICLR 2025 \ Wen Wang, Qiuyu Wang, Kecheng Zheng, Hao Ouyang, Zhekai Chen, Biao Gong, Hao Chen, Yujun Shen, Chunhua Shen
Interactive frame interpolation with flexible drag-based local motion control. Produces diverse transitions from the same start and end frames and supports image morphing.

FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior \ ECCV 2024 \ Zhekai Chen*, Wen Wang*, Zhen Yang, Zeqing Yuan, Hao Chen†, Chunhua Shen†
A generic zero-shot framework for appearance and semantic image composition using diffusion priors. Extends to object removal and multi-character customization without task-specific training.

AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort \ International Journal of Computer Vision \ Wen Wang*, Canyu Zhao*, Hao Chen, Zhekai Chen, Kecheng Zheng, Chunhua Shen
Generates text-aligned, identity-consistent storytelling images from stories and character references. Uses language-model planning and dense condition generation to reduce manual control requirements.
🧭 Journey
Tencent Hunyuan
Qingyun Program · Research Intern
The University of Hong Kong
Ph.D. in ECE · HKU MMLab
Zhejiang University
B.Eng. in Computer Science and Technology · CAD&CG Lab
Suzhou Academy
High School · Key Class
🤝 Collaborations
Wen Wang
A senior colleague who helped me get started in visual generation research and guided my early projects.
Yujie Wei
Generative models, controllable generation, and model scaling
🏆 Awards
- Sep. 2025. HKU Presidential PhD Scholar Programme (HKUPS).
- Dec. 2024. Ho Chi Kwan Education Scholarship, Zhejiang University (7 recipients each year).
- Oct. 2023. Third Prize, Sixth Open Source Innovation Competition.
🛎 Academic Service
- Conference Reviewer: ICLR 2025, ICLR 2026, ECCV 2026.
- Journal Reviewer: IEEE TPAMI, International Journal of Computer Vision (IJCV).





