OpenBot
Egocentric datasetOpenReadiness 70/100Provisional

HumanPlus

HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior…

Source notes

HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior cloning on teleoperated demonstrations. The low-level shadowing policy is trained in simulation using the 40-hour AMASS human-motion dataset; the released task data consists of HDF5 imitation-learning episodes (ACT/Mobile-ALOHA style) recording two head-mounted egocentric RGB cameras plus 19-DoF body and dexterous-hand joint positions. Demonstrated skills include folding clothes, rearranging objects, warehouse unloading, two-robot greeting, wearing a shoe, and typing.

Scale
40 hours
Formats
hdf5
License
unknown
Published
2024-06-15

Decision summary

Best for

teleoperation

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

How scores work

70/100

Provisional

3/6 dimensions scored · 42% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readiness70conf. 55
World-model readinessNot scoredconf. 15
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 70 · confidence 55

World-model readiness

World-model observation, geometry, or temporal semantics are not verified.

unknownnot scored · confidence 15

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Feedback / correction / failure. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.
Engineering checks (OBRS) behind this record
Engineering checksBronze

Needs Audit

4/12checks passed

OBRS metadata and reported test evidence. A breakdown behind Selection readiness, not a separate score or an independent OpenBot certification.

Standardization & Loaders8 / 25 pt
Physical & Action Quality8 / 25 pt
Semantic & Annotation12 / 20 pt
Real-World Validation0 / 15 pt
License & Compliance0 / 15 pt
Missing evidence and readiness gaps
  • A passing real-hardware test report is required.
  • A passing ingestion/pipeline test report is required.
  • A passing privacy and provenance review is required.

Dataset facts

Source
Stanford University
Evidence
secondary claim
Formats
hdf5
hours
40
tasks
6
Source-reported scale
240 trajectories
Read paper

Loop signals

5/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
View 5 more signal categories
Observation / ego video

video · HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior cloning on teleoperated demonstrations. The low-level shadowing policy is trained in simulation using the 40-hour AMASS human-motion dataset; the released task data consists of HDF5 imitation-learning episodes (ACT/Mobile-ALOHA style) recording two head-mounted egocentric RGB cameras plus 19-DoF body and dexterous-hand joint positions. Demonstrated skills include folding clothes, rearranging objects, warehouse unloading, two-robot greeting, wearing a shoe, and typing.

present
Action / hand pose / robot state

HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior cloning on teleoperated demonstrations. The low-level shadowing policy is trained in simulation using the 40-hour AMASS human-motion dataset; the released task data consists of HDF5 imitation-learning episodes (ACT/Mobile-ALOHA style) recording two head-mounted egocentric RGB cameras plus 19-DoF body and dexterous-hand joint positions. Demonstrated skills include folding clothes, rearranging objects, warehouse unloading, two-robot greeting, wearing a shoe, and typing.

partial
Language intent / task phase

HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior cloning on teleoperated demonstrations. The low-level shadowing policy is trained in simulation using the 40-hour AMASS human-motion dataset; the released task data consists of HDF5 imitation-learning episodes (ACT/Mobile-ALOHA style) recording two head-mounted egocentric RGB cameras plus 19-DoF body and dexterous-hand joint positions. Demonstrated skills include folding clothes, rearranging objects, warehouse unloading, two-robot greeting, wearing a shoe, and typing.

partial
Sim-real pairing

HumanPlus is a Stanford full-stack system that lets a customized 33-DoF Unitree H1 humanoid shadow human body and hand motion in real time from RGB cameras, and then learn autonomous whole-body skills via behavior cloning on teleoperated demonstrations. The low-level shadowing policy is trained in simulation using the 40-hour AMASS human-motion dataset; the released task data consists of HDF5 imitation-learning episodes (ACT/Mobile-ALOHA style) recording two head-mounted egocentric RGB cameras plus 19-DoF body and dexterous-hand joint positions. Demonstrated skills include folding clothes, rearranging objects, warehouse unloading, two-robot greeting, wearing a shoe, and typing.

present
License / format / access

Open · unknown · hdf5

partial
Evidence details and provenance

Official signal claims

ProprioceptionVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.