Yue Su

Experience

A continuous path through
research and robotics.

Xidian University

B.Eng. · OMEGA Lab
Rank 4 / 174 · National Scholarship

MVIG · SJTU

Research Assistant
Robot learning

Astribot

Research Intern
Whole-body manipulation

ByteDance Seed

Research Intern
Seed Robotics

MMLab · HKU

PhD
World models & action

Experience02
Selected works

Four projects, one research thread.

Condition-guided action generation03
MBA · 01 / 02

Motion before action.

First imagine the object trajectory—then let it guide the robot trajectory.

Cut clay
Open drawer
Put bread into pot
Pour balls
MBA · 02 / 02

Two diffusion processes, one manipulation condition.

Motion Before Action pipeline
Step 1

Motion diffusion

Predict object motion from the current observation and language instruction.

Step 2

Action diffusion

Condition the manipulation policy on the generated motion representation.

Property

Plug and play

Attach MBA to existing manipulation policies without redesigning the action backbone.

Dense Policy · 01 / 02

Action generation is a design choice.

Comparison of holistic, next-token, next-chunk, and Dense Policy action generation

Sparse first.
Dense later.

Dense Policy predicts a few structural keyframes, then expands them bidirectionally in logarithmic inference steps.

Dense Policy · 02 / 02

Expand from sparse keyframes—in both directions.

Dense Policy pipeline
Bidirectional generation

Sparse → dense.

Predict key actions first, then recursively fill the intervals to form a coherent trajectory.

ICCV 2025 paper ↗
Real-robot behavior

Coherent manipulation.

The policy preserves global task structure while refining detailed local actions.

See all demos ↗
Bidirectional autoregressive learning of actions07
DSPv2 · 01 / 02

Whole-body mobile manipulation changes the problem.

DSPv2 overview
Perception

See what matters

Mobile settings introduce larger visual variation and longer interaction horizons.

Coordination

Move as one system

Base, torso, and arm actions must form a coherent whole-body trajectory.

Generalization

Leave the training setup

A useful policy must transfer across objects, tasks, and environments.

DSPv2 · 02 / 02

Perceive globally. Act coherently.

DSPv2 perception and action pipeline
World Guidance · Motivation
Learn what to predict for action—not what humans predefine.

Raw future pixels are rich but redundant. Hand-designed targets are compact but limiting. WoG learns the bottleneck: a condition space that retains precisely what action needs.

Manual targetsVideo
Depth
Semantics
Hand-crafted objectives are detailed, but limited and difficult to generalize.
World GuidanceAction-useful
condition
Compact, learned from data, and intrinsically aligned with action.
World Guidance · One view

World modeling inside the condition space.

World Guidance method overview
Compact

Less is more

Compress future observations into a non-redundant predictive space.

Predictable

Action-centric latent

The VLA predicts conditions that are intrinsically relevant to action.

Fine-grained

Guidance that acts

Future-aware conditions support precise and generalizable manipulation.

World Guidance · Method

Condition in. Condition out.
Complete inference.

World Guidance two-stage inference pipeline
Stage I

Condition in

Inject future observations into the action inference pipeline.

Stage II

Condition out

Predict compressed future conditions alongside actions.

Transfer

Guidance retained

Carry future knowledge into the VLA without requiring future frames at test time.

World Guidance · Real-world behavior

Fine-grained action generation.4× speed

“put the green cup into the plate”
“put the red cup into the plate” (unseen)
“fold the brown towel” (unseen)
“fold the blue towel”
“fold the white towel”
“close the microwave”
“fold the blue towel” (light change)
“put the green cup into the plate”
“put the green cup into the plate” (background change)
“fold the white towel” (background change)
“put the green cup into the plate”
“fold the white towel”
World Guidance · Human videos

Any video can provide a condition.

Human manipulation trajectories used by World Guidance
1920hhuman manipulation
video
220h11.5% annotated
with actions
Learning from human manipulation videos14

Thank you.

MBADense PolicyDSPv2World Guidance
Yue Su
01 / 15