Hi there
Welcome to my Homepage!
I am a PhD student at
MMLab@HKU (2026.9 - 2030.8), advised by Prof. Xihui Liu. I founded IAAI, and currently intern at Tencent LightSpeed.
I received my B.Eng. from Xidian University (2022.9 - 2026.6). During my studies, I interned at ByteDance Seed and was an RA at MVIG@SJTU with Prof. Lixin Yang and Prof. Cewu Lu.
News
- Internship completed
Experience

The University of Hong Kong Sep 2026 - Phd at MMLab@HKU

Tencent IEG Sep 2026 - Talent Prog. Intern at LightSpeed

ByteDance Seed Oct 2025 - May 2026 Research Intern at Seed-Robotics

Astribot Inc. June 2025 - Sep 2025 Research Intern with Jianan Wang

Shanghai Jiao Tong University July 2024 - June 2025 Research Assistant at MVIG Lab


湖北省武昌实验中学 Sep 2019 - June 2022 那是一段小有遗憾的幸福时光.
Works
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Ruiteng Zhao, Zhengshen Zhang, , Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu†
SG-WAM predicts action-conditioned future dynamics in a geometry-aware policy representation space for robust robot manipulation.
Ruiteng Zhao, Zhengshen Zhang, , Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu†
SG-WAM predicts action-conditioned future dynamics in a geometry-aware policy representation space for robust robot manipulation.

CARA: Concept-Aware Risk Attention for Interpretable Collision Prediction
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, , Jinbo Su, Yi Hong†
An interpretable spatio-temporal framework that grounds collision prediction in evolving, human-understandable risk concepts.
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, , Jinbo Su, Yi Hong†
An interpretable spatio-temporal framework that grounds collision prediction in evolving, human-understandable risk concepts.
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang*, Jianan Wang*, Zifeng Gao, , Tiranru Dai, Cai Zhou,
Jiwen Lu, Yansong Tang†
Learning Vision-Language-Action Models from Human Videos.
Chubin Zhang*, Jianan Wang*, Zifeng Gao, , Tiranru Dai, Cai Zhou,
Jiwen Lu, Yansong Tang†
Learning Vision-Language-Action Models from Human Videos.

AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors
Hao Li†, Fanggao Wan, , Yue Wu, Mingyang Zhang, Maoguo Gong†
Historically, infrared adversarial attacks were single-use and tough to deploy. Using TEC, we implemented efficient attacks adaptable to hardware scenarios.
Hao Li†, Fanggao Wan, , Yue Wu, Mingyang Zhang, Maoguo Gong†
Historically, infrared adversarial attacks were single-use and tough to deploy. Using TEC, we implemented efficient attacks adaptable to hardware scenarios.
REPEAT. REFINE.
Aha Looped Transformer
An open visual atlas exploring looped Transformers, recurrent computation, adaptive depth, and latent reasoning.
An open visual atlas exploring looped Transformers, recurrent computation, adaptive depth, and latent reasoning.
ManiUniCon: A Unified Control Interface for Robotic Manipulation
ManiUniCon is a comprehensive, multi-process robotics control framework designed for robotic manipulation tasks. It provides a unified interface for controlling various robot arms, integrating sensors, and executing policies in real-time.
ManiUniCon is a comprehensive, multi-process robotics control framework designed for robotic manipulation tasks. It provides a unified interface for controlling various robot arms, integrating sensors, and executing policies in real-time.
Novels
Kindred· An AI simulated cityBlogs
Awards
- Xiaomi Outstanding Scholarship
- National Scholarship
- Outstanding Student




