RoboTwin 2.0 Leaderboard

Benchmark Setting

Training data is fixed to 50 demo_clean trajectories × 50 tasks (2,500 demos total) on Aloha-AgileX, then evaluated 100 trials/task under demo_clean and demo_randomized. Use the ranking control to switch titles between Average(c2c+c2r), clean2random(hard), and clean2clean(easy). Co-train and Single-task SFT share one board. The default ranking is Average(c2c+c2r).

Submission & listing policy

To be listed on this leaderboard, a model must provide publicly released code, publicly released weights, and a technical report (arXiv paper or equivalent public document) describing the method. To submit results or request listing, contact chentianxing2002@gmail.com with links to the code, weights, and report.

XPolicyLab

A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

Institutions MMLab@HKU & THU
Contact Tianxing Chen (project lead), chentianxing2002@gmail.com

The same adapters serve this RoboTwin 2.0 board and RoboDojo (sim + real). Results here are reproduced via the XPolicyLab standard interface.

Data: 50 demo_clean × 50 tasks Co-train: one policy jointly trained on all 50 tasks Single: one checkpoint fine-tuned per task (paper baselines) Average(c2c+c2r): ranked by (c2c + c2r) / 2 clean2random(hard): ranked by demo_randomized clean2clean(easy): ranked by demo_clean Open: public code + weights + technical report required XPolicyLab: bridges RoboTwin ↔ RoboDojo Contributor: RoboTwin Team
Ranking setting
View
Ranked by Average(c2c+c2r) mean ↓
Current ranking setting

Average(c2c+c2r)

Train: 50 demo_clean × 50 tasks (2,500) Eval / sort: (c2c + c2r) / 2 Robot: Aloha-AgileX Track: Co-train (+ Single SFT baselines)

Overall Ranking — Average(c2c+c2r)

Latest update:
Rank Method Contributor Average(c2c+c2r) clean2random(hard) clean2clean(easy) Track Date

Per-task Results

Columns follow the current Average(c2c+c2r) / clean2random(hard) / clean2clean(easy) ranking. Methods tagged Single are single-task SFT baselines.

RoboTwin XPolicyLab RoboDojo

XPolicyLab bridges RoboTwin and RoboDojo with one policy interface, data format, and train/eval pipeline. Website · Code

Sim Leaderboard Top 10 · Ranked by Average · Score ↓

Dimension
Rank by

XPolicyLab bridges RoboTwin and RoboDojo: one policy interface, one data format, and one train/eval pipeline. Models listed here can be transferred to RoboDojo (sim + real) without rewriting a line of code. Preview below shows the official Top 10; cells are Score / SR% per capability dimension.

Show full leaderboard on robodojo-benchmark.com

News