OpenBot
Back to Explore
Manipulation datasetOpen

EgoDex

EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

Scale
829 hours
Formats
hdf5
License
research-only
Published
2025-05-16

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

73/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
Apple
Evidence
secondary claim
Formats
hdf5
episodes
338000
hours
829
tasks
194
bytes
2199023255552
Read paper

Metadata coverage

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

video · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

present
Action / hand pose / robot state

ee_pose · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

partial
Language intent / task phase

language · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

present
Feedback / correction / failure

EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

partial
Sim-real pairing

EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.

present
License / format / access

Open · research-only · hdf5

partial
Evidence details and provenance

Official signal claims

Ee_poseLanguageVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records