RoboDojo
RoboDojo is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of 30 policies evaluated in simulation…
Source notes
RoboDojo is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of 30 policies evaluated in simulation reached only 8.80% success against 76.03% for human teleoperation.
- 60 tasks total: 42 simulation tasks (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus 18 real-world tasks across three embodiments (ARX X5, Piper, Piper X).
- Ships a training dataset: 3,500 simulated trajectories (1,859,602 frames, 20.66 h) and 1,800 real-world trajectories (1,611,841 frames, 17.91 h), all bimanual and recorded at 25 Hz.
- Distributed in several formats — LeRobot v3.0 (120 GB), LeRobot v2.1 (64 GB), HDF5 (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB).
- Real-world rigs use 3 synchronised cameras: one head camera (Gemini 335L) and two wrist cameras (Gemini 305). Evaluation is 10 trials × 18 tasks = 180 real trials per policy, scored double-blind by three independent evaluators.
- RoboDojo-RealEval is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any.
- Policy integration runs through XPolicyLab, which unifies 40+ policies behind one interface.
- Simulation demonstrations come from two sources: automated trajectory synthesis composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and VR teleoperation for tasks too complex to synthesise. Real-world demos are leader-follower teleoperation, 100 per task from four operators.
Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
- Scale
- 38.6 hours
- Formats
- hdf5 · lerobot
- License
- Apache-2.0
- Published
- 2026-07-05
Decision summary
scripted
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from official dataset metadata.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
77/100
Provisional
4/6 dimensions scored · 48% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- MMLab@HKU and collaborators
- Evidence
- secondary claim
- Formats
- hdf5 · lerobot
- episodes
- 5,300
- hours
- 38.6
- tasks
- 60
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
depth · video · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
presentee_pose · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
presentlanguage · has-success-labels · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
present**RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
presentdepth · simulation · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.
presentOpen · Apache-2.0 · hdf5 · lerobot
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
