About
I am currently a master student at Beijing University of Posts and Telecommunications, advised by Prof. Kongming Liang and Prof. Zhanyu Ma at the PRIS-CV Lab. I earned my Bachelor's degree in a joint program between Beijing University of Posts and Telecommunications and Queen Mary University of London in 2024.
My current research focuses on Embodied AI. I am currently seeking PhD opportunities for Fall 2027. If you are interested in working with me, please feel free to email me :)
News
One paper accepted to ACMMM 2026 (CCF-A)! 🎉
Hepto-LLaVA accepted to MICCAI 2026 (CCF-B)! 🎉
Joined Simplexity Robotics as a Research Intern, led by Peng Jia 🧑💻
Joined Peking University as a Research Assistant, supervised by Shanghang Zhang 🧑💻
Joined Shanghai AI Lab as a Remote Research Assistant, supervised by Weijia Li and Conghui He 🧑💻
Released Step-GUI! 🎉
MedReasoner accepted to AAAI 2026 (CCF-A)! 🎉
Joined StepFun as a Research Intern, led by Zheng Ge and Xiangyu Zhang 🧑💻
TRIG accepted to ICCV 2025 (CCF-A)! 🎉
PGP-SAM accepted to ISBI 2025! 🎉
Joined BUPT AI, PRIS-CV as a Master Student, advised by Kongming Liang and Zhanyu Ma 👨🎓
Selected as an outstanding graduate of Beijing! 🎉
One paper accepted to MICCAI 2024 (CCF-B)! 🎉
Research Trajectory
Benchmarks for multimodal intelligence, mutually reinforcing reasoning and grounding, and post-training for embodied foundation models.
Evaluation
Developing benchmarks for AIGC systems and MLLMs.
Multimodal Large Language Models
Exploring reasoning and grounding as mutually reinforcing capabilities in MLLMs.
Embodied Foundation Models
Advancing VLA and WAM through foundation-model post-training.
Evaluation
Developing benchmarks for AIGC systems and MLLMs.
Multimodal Large Language Models
Exploring reasoning and grounding as mutually reinforcing capabilities in MLLMs.
Embodied Foundation Models
Advancing VLA and WAM through foundation-model post-training.
Selected Publications
View All →*: Equal Contribution, †: Project Lead, ‡: Corresponding Author

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
Hao Chen*, Jiaming Liu*†, Zhonghao Yan*, Nuowei Han*, Renrui Zhang†, Chenyang Gu, Jialin Gao, Ziyu Guo, Siyuan Qian, Yinxi Wang, Peng Jia, Chi-Wing Fu, Shanghang Zhang‡, Pheng-Ann Heng

MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
Zhonghao Yan*, Muxi Diao*, Yuxuan Yang, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu, Kongming Liang‡, Zhanyu Ma

Trade-offs in image generation: How do different dimensions interact?
Sicheng Zhang*, Binzhu Xie*, Zhonghao Yan*, Yuli Zhang, Donghao Zhou, Xiaofei Chen, Shi Qiu, Jiaqi Liu, Guoyang Xie‡, Zhichao Lu
