OpenBot
Back to Explore
Manipulation datasetOpen

Mobile ALOHA

Mobile ALOHA is a teleoperated mobile bimanual manipulation dataset collected by Fu, Zhao, and Finn (Stanford) using the low-cost whole-body Mobile ALOHA system, with ~50 human demonstrations per task across whole-body tasks such as sauteing shrimp, opening a two-door cabinet, calling/entering an elevator, and rinsing a pan. The TFDS/Open X release contains 276 episodes with 3 RGB cameras (overhead + two wrist cameras at 480x640), a 14-dim state, a 16-dim action, and per-step language instructions. Raw HDF5 demonstrations are also distributed via the project's Google Drive.

Scale
276 episodes
Formats
hdf5 · rlds
License
CC-BY-4.0
Published
2024-01-04

Decision summary

Best for

teleoperation

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

73/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Feedback / correction / failure. Not enough evidence is available to classify this signal.
  • Sim-real pairing. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
Stanford University
Evidence
secondary claim
Formats
hdf5 · rlds
episodes
276
bytes
50917757317
Read paper

Metadata coverage

Loop signals

4/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
Sim-real pairing

No decision-grade evidence captured yet.

unknown
View 4 more signal categories
Observation / ego video

video · Mobile ALOHA is a teleoperated mobile bimanual manipulation dataset collected by Fu, Zhao, and Finn (Stanford) using the low-cost whole-body Mobile ALOHA system, with ~50 human demonstrations per task across whole-body tasks such as sauteing shrimp, opening a two-door cabinet, calling/entering an elevator, and rinsing a pan. The TFDS/Open X release contains 276 episodes with 3 RGB cameras (overhead + two wrist cameras at 480x640), a 14-dim state, a 16-dim action, and per-step language instructions. Raw HDF5 demonstrations are also distributed via the project's Google Drive.

present
Action / hand pose / robot state

Mobile ALOHA is a teleoperated mobile bimanual manipulation dataset collected by Fu, Zhao, and Finn (Stanford) using the low-cost whole-body Mobile ALOHA system, with ~50 human demonstrations per task across whole-body tasks such as sauteing shrimp, opening a two-door cabinet, calling/entering an elevator, and rinsing a pan. The TFDS/Open X release contains 276 episodes with 3 RGB cameras (overhead + two wrist cameras at 480x640), a 14-dim state, a 16-dim action, and per-step language instructions. Raw HDF5 demonstrations are also distributed via the project's Google Drive.

partial
Language intent / task phase

language · Mobile ALOHA is a teleoperated mobile bimanual manipulation dataset collected by Fu, Zhao, and Finn (Stanford) using the low-cost whole-body Mobile ALOHA system, with ~50 human demonstrations per task across whole-body tasks such as sauteing shrimp, opening a two-door cabinet, calling/entering an elevator, and rinsing a pan. The TFDS/Open X release contains 276 episodes with 3 RGB cameras (overhead + two wrist cameras at 480x640), a 14-dim state, a 16-dim action, and per-step language instructions. Raw HDF5 demonstrations are also distributed via the project's Google Drive.

present
License / format / access

Open · CC-BY-4.0 · hdf5 · rlds

present
Evidence details and provenance

Official signal claims

LanguageProprioceptionVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records