Latent Action Control for Reasoning-Guided Unified Image Generation
Makes reasoning actionable inside a unified image generator through role-structured latent actions for planning, drafting, diagnosis, and refinement.
Generative AI · Image Generation · Visual Reasoning
I am an MPhil student in the Data Science and Analytics Thrust, Information Hub, at The Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Lei Zhu.
My research focuses on image generation and generative multimodal models. I am particularly interested in making visual understanding actionable for generation through unified understanding-generation models, latent reasoning and control, and reward/evaluation models for generative systems.
High-quality and controllable generative models for visual content creation.
Shared multimodal models that connect visual understanding with image synthesis.
Planning, diagnosis, and refinement through latent actions inside generative models.
Reliable preference modeling and evaluation for image-generation systems.
Makes reasoning actionable inside a unified image generator through role-structured latent actions for planning, drafting, diagnosis, and refinement.
A self-evolving image-generation agent that distills structured visual experience from tool-orchestrated generation trajectories.
A reward model and evaluation benchmark designed for typography-, layout-, and aesthetics-aware assessment of graphic design generation.
MPhil in Data Science and Analytics · Information Hub
Advisor: Prof. Lei Zhu
B.Eng. in Computer Science and Technology · School of Intelligence Science and Technology (School of Future Technology)
Excellent Graduate · Top 12% in major