RoboTwin 2.0 Leaderboard

clean→random · 50 tasks · 50 demo_clean trajectories / task · Updated

Benchmark Setting

Official protocol: clean2random. Policies are trained on Aloha-AgileX, then evaluated 100 trials/task under demo_clean (Easy) and demo_randomized (Hard). Co-train and Single-task SFT are listed in one board; use the sort control to rank by Easy or Hard.

XPolicyLab
Results on this leaderboard are reproduced via the XPolicyLab standard interface.
Co-train: one policy jointly trained on all 50 tasks Single: one checkpoint fine-tuned per task (paper baselines) Easy = demo_clean  |  Hard = demo_randomized Contributor: RoboTwin Team
Rank by
Sorted by Hard (demo_randomized) mean ↓
Primary track: Co-train
50 demo_clean × 50 tasks joint training (2,500 clean demos total), then clean→random evaluation.
Latest update:

Overall Ranking

Rank Method Contributor Easy Hard

Per-task Results

Columns follow the current Easy/Hard ranking. Methods tagged Single are single-task SFT baselines.

News