OpenBot
Back to Explore
Manipulation datasetOpen

EgoInfinity

**EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

Scale
1150 episodes
Formats
custom
License
custom
Published
2026-06-16

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

77/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
Rice RobotPI Lab
Evidence
secondary claim
Formats
custom
episodes
1150
bytes
5894369280
Read paper

Metadata coverage

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

depth · video · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

present
Action / hand pose / robot state

ee_pose · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

present
Language intent / task phase

language · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

present
Feedback / correction / failure

**EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

partial
Sim-real pairing

depth · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).

present
License / format / access

Open · custom · custom

present
Evidence details and provenance

Official signal claims

DepthEe_poseLanguageSegmentationVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records