Ego4D
Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13…
Source notes
Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
- Scale
- 3,670 hours
- Formats
- custom
- License
- Ego4D License Agreement
- Published
- 2022-10-01
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from official dataset metadata.
Ego research graph
Related research
Reviewed paper relationships connect this source record to training roles. They do not imply a general model-performance claim.
- Paper
Introduces dataset · CVPR · 2022
Ego4D: Around the World in 3,000 Hours of Egocentric Video
Introduces the Ego4D corpus and benchmark suite.
- Paper
Derives training data · NeurIPS · 2022
Egocentric Video-Language Pretraining
Builds EgoClip from narrated Ego4D video for pre-training.
Training effectsLanguage grounded in the wearer’s view - Paper
Uses for pre-training · CVPR · 2023
Learning Video Representations From Large Language Models
Uses Ego4D video for narrator-assisted representation learning.
Training effectsLanguage grounded in the wearer’s view - Paper
Uses for pre-training · ICCV · 2023
EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
Uses Ego4D-derived egocentric video–language supervision.
Training effectsLanguage grounded in the wearer’s view
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
69/100
Provisional
3/6 dimensions scored · 41% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation/action alignment has not been established.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- Meta AI (Facebook AI Research) and academic consortium
- Evidence
- secondary claim
- Formats
- custom
- hours
- 3,670
- tasks
- 5
- size
- 30.0 TB
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
video · Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
presentEgo4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
partialEgo4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
presentlanguage · Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
presentEgo4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.
presentGated · Ego4D License Agreement · custom
partialEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
