OpenBot
Manipulation datasetOpenReadiness 77/100Provisional

RoboDojo

RoboDojo is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of 30 policies evaluated in simulation…

Source notes

RoboDojo is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of 30 policies evaluated in simulation reached only 8.80% success against 76.03% for human teleoperation.

  • 60 tasks total: 42 simulation tasks (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus 18 real-world tasks across three embodiments (ARX X5, Piper, Piper X).
  • Ships a training dataset: 3,500 simulated trajectories (1,859,602 frames, 20.66 h) and 1,800 real-world trajectories (1,611,841 frames, 17.91 h), all bimanual and recorded at 25 Hz.
  • Distributed in several formats — LeRobot v3.0 (120 GB), LeRobot v2.1 (64 GB), HDF5 (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB).
  • Real-world rigs use 3 synchronised cameras: one head camera (Gemini 335L) and two wrist cameras (Gemini 305). Evaluation is 10 trials × 18 tasks = 180 real trials per policy, scored double-blind by three independent evaluators.
  • RoboDojo-RealEval is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any.
  • Policy integration runs through XPolicyLab, which unifies 40+ policies behind one interface.
  • Simulation demonstrations come from two sources: automated trajectory synthesis composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and VR teleoperation for tasks too complex to synthesise. Real-world demos are leader-follower teleoperation, 100 per task from four operators.

Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

Scale
38.6 hours
Formats
hdf5 · lerobot
License
Apache-2.0
Published
2026-07-05

Decision summary

Best for

scripted

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

How scores work

77/100

Provisional

4/6 dimensions scored · 48% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readiness85conf. 55
World-model readiness82conf. 50
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.
Engineering checks (OBRS) behind this record
Engineering checksBronze

Needs Audit

6/12checks passed

OBRS metadata and reported test evidence. A breakdown behind Selection readiness, not a separate score or an independent OpenBot certification.

Standardization & Loaders15 / 25 pt
Physical & Action Quality8 / 25 pt
Semantic & Annotation20 / 20 pt
Real-World Validation0 / 15 pt
License & Compliance10 / 15 pt
Missing evidence and readiness gaps
  • A passing real-hardware test report is required.
  • A passing ingestion/pipeline test report is required.
  • A passing privacy and provenance review is required.

Dataset facts

Source
MMLab@HKU and collaborators
Evidence
secondary claim
Formats
hdf5 · lerobot
episodes
5,300
hours
38.6
tasks
60
Read paper

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

depth · video · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

present
Action / hand pose / robot state

ee_pose · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

present
Language intent / task phase

language · has-success-labels · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

present
Feedback / correction / failure

**RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

present
Sim-real pairing

depth · simulation · **RoboDojo** is a sim-and-real evaluation benchmark built to answer a blunt question: how good are generalist manipulation policies really? The answer it reports is sobering — the best of **30 policies evaluated in simulation** reached only **8.80%** success against **76.03%** for human teleoperation. - **60 tasks total**: **42 simulation tasks** (NVIDIA Isaac Sim 5.1 / Isaac Lab 2.3) spanning five capability axes — Generalization (12), Memory (6), Precision (8), Long-Horizon (8), Open-Vocabulary (8) — plus **18 real-world tasks** across three embodiments (ARX X5, Piper, Piper X). - Ships a **training dataset**: **3,500 simulated trajectories** (1,859,602 frames, **20.66 h**) and **1,800 real-world trajectories** (1,611,841 frames, **17.91 h**), all bimanual and recorded at **25 Hz**. - Distributed in several formats — **LeRobot v3.0** (120 GB), LeRobot v2.1 (64 GB), **HDF5** (523 GB), depth HDF5 (~4.5 TB), and real-world data (273 GB). - Real-world rigs use **3 synchronised cameras**: one head camera (Gemini 335L) and **two wrist cameras** (Gemini 305). Evaluation is **10 trials × 18 tasks = 180 real trials** per policy, scored double-blind by three independent evaluators. - **RoboDojo-RealEval** is a cloud-accessible real-robot evaluation service with standardised hardware and scene reset, so a policy can be benchmarked on physical robots remotely without owning any. - Policy integration runs through **XPolicyLab**, which unifies **40+ policies** behind one interface. - Simulation demonstrations come from two sources: **automated trajectory synthesis** composed from motion-planning primitives (grasp, place, handover, insert, open, close, stack, push_up), and **VR teleoperation** for tasks too complex to synthesise. Real-world demos are **leader-follower teleoperation**, 100 per task from four operators. Assembled by a 43-author consortium led by MMLab@HKU, with contributors from UC Berkeley, Tsinghua, Peking University, Stanford and MIT.

present
License / format / access

Open · Apache-2.0 · hdf5 · lerobot

present
Evidence details and provenance

Official signal claims

DepthEe_poseLanguageProprioceptionVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.