EgoInfinity
**EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
- Scale
- 1150 episodes
- Formats
- custom
- License
- custom
- Published
- 2026-06-16
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
77/100
Provisional · confidence 48 · 4/6 evaluated
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Record specifics
Dataset facts
- Source
- Rice RobotPI Lab
- Evidence
- secondary claim
- Formats
- custom
- episodes
- 1150
- bytes
- 5894369280
Metadata coverage
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
depth · video · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
presentee_pose · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
presentlanguage · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
present**EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
partialdepth · **EgoInfinity** is not a static dataset but a **modular, upgradable data engine**: it takes ordinary internet video of people manipulating objects and automatically recovers metric 4D hand–object interaction, then retargets that motion into executable robot joint trajectories — no manual annotation and no depth sensors. - Runs over Meta FAIR's **Action100M** corpus — **14.6 years (~127,000 h)** of footage across **~142M action clips**. The publicly released artifact is a curated **preview subset of 1,150 rows / 104 fully retargeted clips (5.49 GB)**. - **Eight-stage pipeline**: metric depth + focal estimation (MoGe-2), gravity calibration (GeoCalib), hand detection and MANO reconstruction (YOLO + WiLoR), optical flow and depth stabilisation (MEMFOF), motion infilling (HaWoR), text-prompted detection and tracking (SAM 3.1 + SAM 2), single-image 3D object reconstruction, and 6-DoF mesh pose tracking. - Each processed clip yields per-frame **21-joint hand poses**, MANO meshes, **per-object 6-DoF pose trajectories**, Gaussian-splat object meshes, optical flow, metric depth, and segmentation masks. - **Retargets to four embodiments**: Franka Panda (7-DoF + gripper), Unitree G1 (7-DoF arm + 7-DoF hand), NASA Robonaut 2 (7-DoF arm + 18-DoF hand), and XLeRobot (5-DoF + gripper), with per-frame IK-success and manipulability metrics stored alongside. - Supports **arbitrary camera viewpoints and shot sizes**, including partial body visibility — unlike prior pipelines that assume a fixed egocentric rig. - Validated on a real robot for **grasping, cutting, wiping, and pouring** learned directly from internet video. - Because it is an engine rather than a fixed corpus, data quality improves automatically as its component models are swapped for better ones. Licensing is **non-commercial research only** (FAIR Noncommercial Research License v1 for data; the code carries mixed MIT / CC-BY-NC-ND / AGPL components).
presentOpen · custom · custom
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
