HiFi-UMI-2K
HiFi-UMI-2K is the public release from HiFi-UMI, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper…
Source notes
HiFi-UMI-2K is the public release from HiFi-UMI, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop.
The claim being tested. Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the fidelity of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA.
Capture. Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs.
Six views per episode. A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps.
What is in the release. 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. observation.state and action are both float32[20] — per hand, [x, y, z, rot6d ×6, gripper_angle_rad] for right then left. The action is an absolute next-state target, not a relative delta, and the authors warn against assuming otherwise.
Validity masks, not success labels. Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; valid.frame == true is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only.
The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
- Scale
- 2,137 hours
- Formats
- lerobot
- License
- CC-BY-4.0
- Published
- 2026
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
No source-backed day-level release date is published for this dataset. Its year is not being converted into an inferred event.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
How scores work73/100
Provisional
4/6 dimensions scored · 48% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Engineering checks (OBRS) behind this record
Needs Audit
OBRS metadata and reported test evidence. A breakdown behind Selection readiness, not a separate score or an independent OpenBot certification.
- A passing real-hardware test report is required.
- A passing ingestion/pipeline test report is required.
- A passing privacy and provenance review is required.
Dataset facts
- Source
- Simple AI (Simple World Lab)
- Evidence
- secondary claim
- Formats
- lerobot
- episodes
- 482.1K
- hours
- 2,137
- tasks
- 588
Loop signals
7/7 present or partial
video · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentee_pose · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentView 5 more signal categories
HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
partiallanguage · HiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentHiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0.
presentHiFi-UMI-2K is the public release from **HiFi-UMI**, a portable bimanual capture system built by Simple AI's Simple World Lab. Demonstrations are recorded by a human wearing a head-mounted stereo rig and a glove-like gripper on each hand — there is no robot, no teleoperation rig, and no instrumented room anywhere in the collection loop. **The claim being tested.** Robot-free UMI data scales easily but is normally used only for pre-training, with a small amount of real-robot teleoperation added at post-training as an accuracy "anchor". The accompanying paper asks whether raising the *fidelity* of the robot-free data can remove that anchor entirely, and reports that a policy post-trained solely on HiFi-UMI deploys directly on a real bimanual robot — success-rate differences of −2.5, +3.1 and −0.6 points against in-domain teleoperation across StarVLA-QwenPI, OpenPI-π₀.₅ and LingBot-VA. **Capture.** Hand pose comes from offline stereo-inertial SLAM plus fiducial marker cubes localised in the head-camera frame, giving roughly 3 mm workspace-local end-effector accuracy without external tracking. Both hands are solved in the same frame, so inter-hand relative pose is native rather than reconstructed across cameras. A shared GPIO trigger keeps every camera, IMU and gripper signal within 40 µs. **Six views per episode.** A head-mounted stereo pair plus two non-parallel fisheye cameras on each hand, covering about 200° around each gripper — all six at 640×512, h264, 25 fps. **What is in the release.** 482,060 episodes and 192,297,515 frames across 398 independently-loadable LeRobot v3 parts, totalling 16.0 TB. 588 distinct natural-language task strings, drawn from ordinary indoor work: microwaves and rice cookers, drawers and cabinets, whiteboards and meeting-room doors, wiping and tidying. `observation.state` and `action` are both float32[20] — per hand, `[x, y, z, rot6d ×6, gripper_angle_rad]` for right then left. The action is an **absolute next-state target**, not a relative delta, and the authors warn against assuming otherwise. **Validity masks, not success labels.** Frames that fail quality control are deliberately retained so that Parquet rows, video frames and timestamps stay index-aligned; `valid.frame == true` is the intended training filter. There is no per-episode success or failure annotation in the exported metadata — the corpus is curated demonstrations that passed replay validation, so it is treated here as success-only. The 2,000-hour public release is a curated subset of a larger internal corpus the paper describes as exceeding 20,000 hours. Released under CC BY 4.0. · Simple AI (Simple World Lab)
presentOpen · CC-BY-4.0 · lerobot
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
