Hi there Hi Welcome to my Homepage!

I am an incoming Phd at MMLab@HKU (2026.9 - 2030.8) with Prof. Xihui Liu.

Previously I worked at ByteDance Seed, MVIG@SJTU with Prof. Lixin Yang and Prof. Cewu Lu.

I got my B. Eng. degree from Xidian University (2022.9 - 2026.6).

News

  • 2026/05 Finished my internship at ByteDance Seed.

Experience

The University of Hong Kong
Sep 2026 -
Phd at MMLab@HKU
ByteDance Seed
Oct 2025 - May 2026
MLE Intern at Seed-Robotics
Astribot Inc.
June 2025 - Sep 2025
MLE Intern advised by Jianan Wang
Shanghai Jiao Tong University
July 2024 - June 2025
Research Assistant at MVIG Lab
Xidian University
Sep 2022 - July 2026
Rank 4/174, National Scholarship
B.E at SAI & RA at OMEGA Lab
Hubei Wuchang Experimental High School
Sep 2019 - June 2022
那是一段小有遗憾的幸福时光.

Publications

CARA framework overview
CARA: Concept-Aware Risk Attention for Interpretable Collision Prediction
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du, Yilei Yuan, Sherwin Ho, Yue Su, Jinbo Su, Yi Hong
An interpretable spatio-temporal framework that grounds collision prediction in evolving, human-understandable risk concepts.
ACM MM 2026 Oral [arXiv]
Game multiverse survey
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
Kuan Zhang*, Dongchen Liu*, Qiyue Zhao*, Tianyu Xin*, Yue Su*, Haisheng Wang, Han Yin, Hongbo Ma, Peize Li, ..., Yiming Li
A survey of foundation models as generalist game players across datasets, models, harnesses, and benchmarks.
ArXiv Preprint   [arXiv] [code]
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang*, Jianan Wang*, Zifeng Gao, Yue Su, Tiranru Dai, Cai Zhou,
Jiwen Lu, Yansong Tang

Learning Vision-Language-Action Models from Human Videos.
ArXiv Preprint   [机器之心] [arXiv] [code] [website]
DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
Yue Su, Chubin Zhang, Sijin Chen, Liufan Tan,
Yansong Tang, Jianan Wang, Xihui Liu

Improved Dense Policy for Whole-body Mobile Manipulation, with effective perception, generalizable manipulation and coherent actions.
ICRA 2026   [arXiv] [code] [website]
Raa
AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors
Hao Li†, Fanggao Wan, Yue Su, Yue Wu, Mingyang Zhang, Maoguo Gong
Historically, infrared adversarial attacks were single-use and tough to deploy. Using TEC, we implemented efficient attacks adaptable to hardware scenarios.
AAAI 2025   [paper]
Dense Policy: Bidirectional Autoregressive Learning of Actions
Yue Su*, Xinyu Zhan*, Hongjie Fang, Han Xue,
Haoshu Fang, Yong-Lu Li, Cewu Lu, Lixin Yang

Propose Dense Policy, A bidirectional robotic autoregressive policy, which infers trajectories by gradually expanding actions from sparse keyframes, demonstrated exceeding diffusion policies.
ICCV 2025   [paper] [arXiv] [website] [3D-code] [2D-code]
MBA
Motion Before Action: Diffusing Object Motion as Manipulation Condition
Yue Su*, Xinyu Zhan*, Hongjie Fang, Yong-Lu Li, Cewu Lu, Lixin Yang
Propose MBA, a novel plug-and-play module leveraging cascaded diffusion processes to generate actions guided by object motion, enabling seamless integration with manipulation policies.
RA-L 2025, ICRA 2026  [paper] [arxiv] [website] [code]
RIaa
Generative Adversarial Patches for Physical Attacks on Cross-Modal Pedestrian Re-Identification
Yue Su, Hao Li†, Maoguo Gong
A generative physical adversarial attack on VI-ReID models perturbs modality-invariant features.
ArXiv Preprint   [arxiv]

Projects

Maniunicon
ManiUniCon: A Unified Control Interface for Robotic Manipulation
ManiUniCon is a comprehensive, multi-process robotics control framework designed for robotic manipulation tasks. It provides a unified interface for controlling various robot arms, integrating sensors, and executing policies in real-time.
Universal-Control Team   [code]
MetaPalace
MetaPalace: Let you in a meta world of The Palace Museum
We've done what the Old Palace official website couldn't: offering 3D artifact views with single-view reconstruction and an interactive LLM-powered tour guider using RAG technology.
[website] [front-end code] [back-end code]

Awards

  • 2025 Xiaomi Outstanding Scholarship
  • 2025 National Scholarship
  • 2025 Outstanding Student, Xidian University

Talks