I am a Ph.D. student in Electrical and Computer Engineering at The University of Hong Kong, advised by Prof. Xihui Liu at HKU MMLab. I received my B.Eng. from Zhejiang University, where I worked with Prof. Chunhua Shen and Prof. Hao Chen at the State Key Laboratory of CAD&CG.

My research interests lie in generative models, unified multimodal models, and agents. I am particularly interested in scalable visual generation, long-context multimodal learning, and generalist agents that can perceive, reason, and act in real-world environments.

I am currently a research intern in the Tencent Hunyuan Qingyun Program. Previously, I worked at Meituan with Manyuan Zhang as my mentor, and at Alibaba Tongyi Wan, where Ruihang Chu mentored me on video generation research.

🔥 News

  • Jul. 2026.  🤖 We release UniClawBench, a universal benchmark with 400 bilingual real-world tasks for proactive agents.
  • Jul. 2026.  💼 I join the Tencent Hunyuan Qingyun Program as a research intern.
  • Mar. 2026.  🎨 We release MACRO, featuring MacroData, a 400K structured long-context dataset for multi-reference image generation.
  • Dec. 2025.  🎬 We release Wan-Move, a scalable framework for precise motion-controllable video generation.
  • Sep. 2025.  🎓 I start my Ph.D. in ECE at The University of Hong Kong and receive the HKU Presidential PhD Scholar Programme award.
  • 2025.  🎉 TTS-VAR and Wan-Move are published at NeurIPS 2025.
  • 2025.  🎉 Framer is published at ICLR 2025.
  • Jun. 2025.  🎓 I receive my B.Eng. from Zhejiang University.
  • 2024.  🎉 FreeCompose is published at ECCV 2024.

📈 Citations

total citations
 h-index ·  i10-index
Google Scholar

Updated from Google Scholar · July 23, 2026

📝 Selected Publications

Selected publications are listed in reverse chronological order by first public submission. Please see Google Scholar for the complete list. * equal contribution  ·  † corresponding author  ·  ‡ project lead

Preprint 2026
UniClawBench overview

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks \ Preprint, 2026 \ Zhekai Chen*, Chengqi Duan*, Kaiyue Sun*, Bohao Li, Yuqing Wang, Manyuan Zhang, Xihui Liu

project | arXiv | github

A capability-driven benchmark containing 400 bilingual, real-world tasks for proactive agents. Evaluates skill usage, exploration, long-context reasoning, multimodal understanding, and cross-platform coordination.

Preprint 2026
MACRO dataset overview

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data \ Preprint, 2026 \ Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu

project | arXiv | github

Introduces MacroData, 400K structured samples with as many as ten reference images. Covers customization, illustration, spatial reasoning, and temporal prediction under long multimodal context.

NeurIPS 2025
Wan-Move motion control examples

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance \ NeurIPS 2025 \ Ruihang Chu*†‡, Yefei He*, Zhekai Chen*, Shiwei Zhang, Xiaogang Xu, Bin Xia, Dingdong Wang, Hongwei Yi, Xihui Liu, Hengshuang Zhao, Yu Liu, Yingya Zhang, Yujiu Yang

project | arXiv | github

Enables precise object and camera motion control using latent trajectory guidance. Integrates motion-aware conditions into an off-the-shelf image-to-video model without architectural changes.

NeurIPS 2025
TTS-VAR framework

TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation \ NeurIPS 2025 \ Zhekai Chen, Ruihang Chu, Yukang Chen, Shiwei Zhang, Yujie Wei, Yingya Zhang, Xihui Liu

arXiv | github

The first general test-time scaling framework for visual autoregressive models. Formulates generation as path search with diversity exploration and potential-based selection.

ICLR 2025
Framer interpolation examples

Framer: Interactive Frame Interpolation \ ICLR 2025 \ Wen Wang, Qiuyu Wang, Kecheng Zheng, Hao Ouyang, Zhekai Chen, Biao Gong, Hao Chen, Yujun Shen, Chunhua Shen

project | arXiv | github

Interactive frame interpolation with flexible drag-based local motion control. Produces diverse transitions from the same start and end frames and supports image morphing.

ECCV 2024
FreeCompose applications

FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior \ ECCV 2024 \ Zhekai Chen*, Wen Wang*, Zhen Yang, Zeqing Yuan, Hao Chen, Chunhua Shen

arXiv | github

A generic zero-shot framework for appearance and semantic image composition using diffusion priors. Extends to object removal and multi-character customization without task-specific training.

IJCV
AutoStory generation examples

AutoStory: Generating Diverse Storytelling Images with Minimal Human Effort \ International Journal of Computer Vision \ Wen Wang*, Canyu Zhao*, Hao Chen, Zhekai Chen, Kecheng Zheng, Chunhua Shen

project | arXiv | github

Generates text-aligned, identity-consistent storytelling images from stories and character references. Uses language-model planning and dense condition generation to reduce manual control requirements.

Other work Image Textualization NeurIPS D&B 2024 Routing Matters in MoE ICLR 2026 Speculative Jacobi-Denoising NeurIPS 2025 DreamVideo-Omni Preprint MSAVBench Preprint

🧭 Journey

Education
Present
Internships
Now
Tencent Hunyuan logo
Jul. 2026 – Present

Tencent Hunyuan

Qingyun Program · Research Intern

HKU logo
Sep. 2025 – Present

The University of Hong Kong

Ph.D. in ECE · HKU MMLab

Now
2026
Meituan logo
Nov. 2025 – Jun. 2026

Meituan

Unified multimodal models · Mentor: Manyuan Zhang

2025
Wan logo
Nov. 2024 – Aug. 2025

Alibaba Tongyi Wan

Video generation · Mentor: Ruihang Chu

Zhejiang University logo
Sep. 2021 – Jun. 2025

Zhejiang University

B.Eng. in Computer Science and Technology · CAD&CG Lab

2025
Suzhou Academy logo
Sep. 2018 – Jun. 2021

Suzhou Academy

High School · Key Class

2021

🤝 Collaborations

Open to collaboration. I enjoy working on generative models, unified multimodal models, and generalist agents. Please feel free to reach out.

🏆 Awards

  • Sep. 2025. HKU Presidential PhD Scholar Programme (HKUPS).
  • Dec. 2024. Ho Chi Kwan Education Scholarship, Zhejiang University (7 recipients each year).
  • Oct. 2023. Third Prize, Sixth Open Source Innovation Competition.

🛎 Academic Service

  • Conference Reviewer: ICLR 2025, ICLR 2026, ECCV 2026.
  • Journal Reviewer: IEEE TPAMI, International Journal of Computer Vision (IJCV).