Ego-Exo4D
Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria…
Source notes
Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
- Scale
- 1,286.3 hours
- Formats
- custom
- License
- Ego-Exo4D License Agreement
- Published
- 2023-11-30
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from official dataset metadata.
Ego research graph
Related research
Reviewed paper relationships connect this source record to training roles. They do not imply a general model-performance claim.
- Paper
Introduces dataset · CVPR · 2024
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Introduces the synchronized Ego-Exo4D corpus and benchmarks.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
69/100
Provisional
3/6 dimensions scored · 41% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation/action alignment has not been established.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- Meta FAIR (Project Aria) and 15+ university partners
- Evidence
- secondary claim
- Formats
- custom
- episodes
- 5,035
- hours
- 1,286.3
- tasks
- 4
Loop signals
7/7 present or partial
video · Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
presentEgo-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
partialView 5 more signal categories
Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
presentlanguage · Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
presentEgo-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
presentEgo-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.
presentGated · Ego-Exo4D License Agreement · custom
partialEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
