Hi there
Welcome to my Homepage!
I am an incoming Phd at MMLab@HKU (2026.9 - 2030.8) with Prof. Xihui Liu.
Previously I worked at ByteDance Seed, MVIG@SJTU with Prof. Lixin Yang and Prof. Cewu Lu.
I got my B. Eng. degree from Xidian University (2022.9 - 2026.6).
News
- Internship completed
Experience






湖北省武昌实验中学
Sep 2019 - June 2022
那是一段小有遗憾的幸福时光.
Sep 2019 - June 2022
那是一段小有遗憾的幸福时光.
Publications
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
Ruiteng Zhao, Zhengshen Zhang, , Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu†
SG-WAM predicts action-conditioned future dynamics in a geometry-aware policy representation space for robust robot manipulation.
Ruiteng Zhao, Zhengshen Zhang, , Wenshuo Wang, Jiahui Li, Zhiyuan Yang, Francis E. H. Tay, Marcelo H. Ang Jr., Haiyue Zhu†
SG-WAM predicts action-conditioned future dynamics in a geometry-aware policy representation space for robust robot manipulation.

CARA: Concept-Aware Risk Attention for Interpretable Collision Prediction
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, , Jinbo Su, Yi Hong†
An interpretable spatio-temporal framework that grounds collision prediction in evolving, human-understandable risk concepts.
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, , Jinbo Su, Yi Hong†
An interpretable spatio-temporal framework that grounds collision prediction in evolving, human-understandable risk concepts.
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang*, Jianan Wang*, Zifeng Gao, , Tiranru Dai, Cai Zhou,
Jiwen Lu, Yansong Tang†
Learning Vision-Language-Action Models from Human Videos.
ArXiv Preprint [机器之心] [arXiv] [code] [website]
Chubin Zhang*, Jianan Wang*, Zifeng Gao, , Tiranru Dai, Cai Zhou,
Jiwen Lu, Yansong Tang†
Learning Vision-Language-Action Models from Human Videos.
ArXiv Preprint [机器之心] [arXiv] [code] [website]

AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors
Hao Li†, Fanggao Wan, , Yue Wu, Mingyang Zhang, Maoguo Gong†
Historically, infrared adversarial attacks were single-use and tough to deploy. Using TEC, we implemented efficient attacks adaptable to hardware scenarios.
AAAI 2025 [paper]
Hao Li†, Fanggao Wan, , Yue Wu, Mingyang Zhang, Maoguo Gong†
Historically, infrared adversarial attacks were single-use and tough to deploy. Using TEC, we implemented efficient attacks adaptable to hardware scenarios.
AAAI 2025 [paper]
Projects
ManiUniCon: A Unified Control Interface for Robotic Manipulation
ManiUniCon is a comprehensive, multi-process robotics control framework designed for robotic manipulation tasks. It provides a unified interface for controlling various robot arms, integrating sensors, and executing policies in real-time.
[team] [code]
ManiUniCon is a comprehensive, multi-process robotics control framework designed for robotic manipulation tasks. It provides a unified interface for controlling various robot arms, integrating sensors, and executing policies in real-time.
[team] [code]

MetaPalace: Let you in a meta world of The Palace Museum
We've done what the Old Palace official website couldn't: offering 3D artifact views with single-view reconstruction and an interactive LLM-powered tour guider using RAG technology.
[website] [front-end code] [back-end code]
We've done what the Old Palace official website couldn't: offering 3D artifact views with single-view reconstruction and an interactive LLM-powered tour guider using RAG technology.
[website] [front-end code] [back-end code]
Awards
- Xiaomi Outstanding Scholarship
- National Scholarship
- Outstanding Student




