Personal introduction · 2026

Yue Su

Learning world models that guide action.
MMLab @ HKU Robot learning World models
Personal · Experience · Researchselen-suyue.github.io ↗
Personal introduction

I build systems that connect imagination to action.

Robot learning,
condition modeling,
and action generation.

My research asks a simple question: what should an agent predict so that its prediction is genuinely useful for control?

VLAWorld ModelsManipulation
Now
Incoming PhD at MMLab, HKU with Prof. Xihui Liu
Before
ByteDance Seed Robotics · Astribot · MVIG, SJTU
Education
B.Eng., Xidian University · Rank 4 / 174
Focus
Action-centric representations that are compact, predictable, and precise
Yue Su · Introduction02
Experience

Across academia and industry,
one continuous research thread.

Xidian University

B.Eng. · OMEGA Lab
National Scholarship

MVIG · SJTU

Research Assistant
Robot learning

Astribot

Research Intern
Whole-body manipulation

ByteDance Seed

Research Intern
Seed Robotics

MMLab · HKU

PhD
World models & action

Experience03
Research

From motion conditions to world guidance.

A progression from structured action generation toward world modeling inside an action-useful condition space.

01 · MBA

Motion as condition

Generate actions guided by future object motion.

Project ↗
02 · Dense Policy

Structure the trajectory

Expand actions bidirectionally from sparse keyframes.

Project ↗
03 · DSPv2

Scale to the whole body

Improve perception, generalization, and action coherence.

Project ↗
04 · World Guidance
Current focus

Model the world in condition space

Learn compact future guidance specifically for action generation.

Explore WoG ↗
Research progression04
MBA · 01 / 02

Motion before action.

Before predicting a robot trajectory, first imagine how the object should move.

Motion Before Action animation
Observation

Objects define the task

The same robot motion can be right or wrong depending on the desired object motion.

Condition

Make the future explicit

Generate a dense object-motion representation before synthesizing the action.

Thesis

Imagination can guide control

A predictive intermediate signal makes manipulation more purposeful.

MBA · 02 / 02

Two diffusion processes, one manipulation condition.

Motion Before Action pipeline
Step 1

Motion diffusion

Predict object motion from the current observation and language instruction.

Step 2

Action diffusion

Condition the manipulation policy on the generated motion representation.

Property

Plug and play

Attach MBA to existing manipulation policies without redesigning the action backbone.

Dense Policy · 01 / 02

Action does not have to be generated left-to-right.

Dense Policy teaser
Problem

Long trajectories are hard

Small early errors compound when every action depends only on its predecessor.

Insight

Keyframes carry structure

Begin with sparse, informative moments that anchor the overall manipulation.

Reframe

Order is a design choice

Autoregressive generation can follow task structure rather than time alone.

Dense Policy · 02 / 02

Expand from sparse keyframes—in both directions.

Dense Policy pipeline
Bidirectional generation

Sparse → dense.

Predict key actions first, then recursively fill the intervals to form a coherent trajectory.

ICCV 2025 paper ↗
Real-robot behavior

Coherent manipulation.

The policy preserves global task structure while refining detailed local actions.

See all demos ↗
Bidirectional autoregressive learning of actions08
DSPv2 · 01 / 02

Whole-body mobile manipulation changes the problem.

DSPv2 overview
Perception

See what matters

Mobile settings introduce larger visual variation and longer interaction horizons.

Coordination

Move as one system

Base, torso, and arm actions must form a coherent whole-body trajectory.

Generalization

Leave the training setup

A useful policy must transfer across objects, tasks, and environments.

DSPv2 · 02 / 02

Effective perception. Generalizable manipulation. Coherent action.

Mobile pick & place
whole-body coordination
Cart manipulation
long-horizon behavior
Object sorting
task diversity
Bowling setup
environment diversity
Improved Dense Policy for whole-body manipulation10
World Guidance · Motivation
Learn what to predict for action—not what humans predefine.

Raw future pixels are rich but redundant. Hand-designed targets are compact but limiting. WoG learns the bottleneck: a condition space that retains precisely what action needs.

Future observationsDetailed, high-dimensional, difficult to predict
WoG
Action-useful conditionsCompact · predictable · fine-grained
World Guidance · One view

World modeling inside the condition space.

World Guidance method overview
Compact

Less is more

Compress future observations into a non-redundant predictive space.

Predictable

Action-centric latent

The VLA predicts conditions that are intrinsically relevant to action.

Fine-grained

Guidance that acts

Future-aware conditions support precise and generalizable manipulation.

World Guidance · Method

Condition in. Condition out.
Complete inference.

World Guidance two-stage inference pipeline
Stage I

Condition in

Inject future observations into the action inference pipeline.

Stage II

Condition out

Predict compressed future conditions alongside actions.

Transfer

Guidance retained

Carry future knowledge into the VLA without requiring future frames at test time.

World Guidance · Real-world behavior

Fine-grained action, even when the world changes.

Put the green cup into the plate
in-distribution
Put the red cup into the plate
unseen object
Fold the blue towel
deformable object
Fold the blue towel
lighting change
Robust control across rigid, articulated, and deformable objects14
World Guidance · Scale & impact

Any video can provide a condition.

1920hhuman manipulation video
220h11% annotated with actions
1

Learn from unannotated video

Condition supervision turns abundant human video into useful training signal.

2

Generalize beyond the scene

The compact condition space filters visual noise while retaining action-relevant dynamics.

3

Connect to real control stacks

Predicted actions can be executed through practical robot interfaces.

World modeling that scales with video15
World Guidance · Takeaways

A smaller prediction space can enable a larger action space.

Modeling is useful only when it changes what the agent can do.

WoG reframes world modeling from visual reconstruction into action-centric guidance.

1

Representation

Learn a compact condition rather than hand-designing the prediction target.

2

Inference

Predict future conditions together with actions—future frames are not needed at deployment.

3

Data

Use abundant human video even when action annotations are sparse.

Practical aside: ManiUniCon ↗ offers a unified real-time manipulation control interface for executing policies across robot setups.
WoG in three points16

Model less.
Guide more.

World Guidance models compact future conditions so an agent can generate precise actions—not merely plausible pixels.

Yue Su · Thank youBack to homepage ↗
YUE SU · SLIDES
01 / 17