OpenBot
Back to Explore
Manipulation datasetOpen

HiFi-UMI-2K

HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

Scale
2137 hours
Formats
lerobot
License
CC-BY-4.0
Published
2026

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

73/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

    Next checks

    • Read the dataset manifest and feature schema.
    • Run a bounded sample audit before assigning Strong readiness.

    Record specifics

    Dataset facts

    Source
    Simple AI (Simple World Lab)
    Evidence
    secondary claim
    Formats
    lerobot
    episodes
    482060
    hours
    2137
    tasks
    588
    bytes
    16042058644025
    Read paper

    Metadata coverage

    Loop signals

    7/7 present or partial

    Observation / ego video

    video · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

    present
    Action / hand pose / robot state

    ee_pose · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

    present
    Gaze / attention

    HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

    partial
    View 4 more signal categories
    Language intent / task phase

    language · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

    present
    Feedback / correction / failure

    HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.

    present
    Sim-real pairing

    HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0. · Simple AI (Simple World Lab)

    present
    License / format / access

    Open · CC-BY-4.0 · lerobot

    present
    Evidence details and provenance

    Official signal claims

    Ee_poseLanguageProprioceptionVideo

    Schema and annotations

    No machine-readable schema facts are captured.

    Sample verification

    Pending. Metadata does not prove sample coverage, alignment, or file integrity.

    curated official source

    Integration notes

    • Metadata requires review against the official source before publication.

    Catalog links

    Related records