EgoDex
EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (90M frames) recorded with Apple Vision Pro.
Source notes
EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
- Scale
- 829 hours
- Formats
- hdf5
- License
- CC-BY-NC-ND-4.0
- Published
- 2025-05-16
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from official dataset metadata.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
73/100
Provisional
4/6 dimensions scored · 48% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
video · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
presentee_pose · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
partiallanguage · EgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
presentEgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
partialEgoDex is a large-scale egocentric human manipulation dataset from Apple, comprising 829 hours of 1080p 30Hz video (~90M frames) recorded with Apple Vision Pro. It pairs each frame with 3D pose annotations for the head, upper body, and hands (68 joints) via on-device tracking, plus camera intrinsics and natural-language task descriptions. The data spans 194 diverse tabletop tasks with everyday objects (~338K episodes), intended for imitation learning of dexterous manipulation from human video.
presentOpen · CC-BY-NC-ND-4.0 · hdf5
partialEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
