HiFi-UMI-2K
HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
- Scale
- 2137 hours
- Formats
- lerobot
- License
- CC-BY-4.0
- Published
- 2026
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
73/100
Provisional · confidence 48 · 4/6 evaluated
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Record specifics
Dataset facts
- Source
- Simple AI (Simple World Lab)
- Evidence
- secondary claim
- Formats
- lerobot
- episodes
- 482060
- hours
- 2137
- tasks
- 588
- bytes
- 16042058644025
Metadata coverage
Loop signals
7/7 present or partial
video · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentee_pose · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentHiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
partialView 4 more signal categories
language · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentHiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentHiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0. · Simple AI (Simple World Lab)
presentOpen · CC-BY-4.0 · lerobot
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
