OpenBot
Egocentric datasetGated
OBRS 36Bronze

Ego4D

Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13…

Source notes

Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

Scale
3,670 hours
Formats
custom
License
Ego4D License Agreement
Published
2022-10-01

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Ego research graph

Related research

Reviewed paper relationships connect this source record to training roles. They do not imply a general model-performance claim.

  1. Introduces dataset · CVPR · 2022

    Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Paper

    Introduces the Ego4D corpus and benchmark suite.

  2. Derives training data · NeurIPS · 2022

    Egocentric Video-Language Pretraining

    Paper

    Builds EgoClip from narrated Ego4D video for pre-training.

  3. Uses for pre-training · CVPR · 2023

    Learning Video Representations From Large Language Models

    Paper

    Uses Ego4D video for narrator-assisted representation learning.

  4. Uses for pre-training · ICCV · 2023

    EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

    Paper

    Uses Ego4D-derived egocentric video–language supervision.

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

69/100

Provisional

3/6 dimensions scored · 41% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readinessNot scoredconf. 15
World-model readiness68conf. 50
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation/action alignment has not been established.

unknownnot scored · confidence 15

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Feedback / correction / failure. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Dataset facts

Source
Meta AI (Facebook AI Research) and academic consortium
Evidence
secondary claim
Formats
custom
hours
3,670
tasks
5
size
30.0 TB
Read paper

Loop signals

6/7 present or partial

Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

video · Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

present
Action / hand pose / robot state

Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

partial
Gaze / attention

Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

present
Language intent / task phase

language · Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

present
Sim-real pairing

Ego4D is a massive-scale dataset of 3,670 hours of daily-life egocentric (first-person) video captured by 923 unique camera wearers across 74 locations in 9 countries, built by a consortium of Facebook AI (Meta) and 13 universities. Portions include audio, 3D environment meshes, eye gaze, stereo, multi-camera footage, IMU, and dense textual narrations, supporting five benchmark suites (episodic memory, hands-and-objects, audio-visual diarization, social interaction, forecasting). It captures human activity only and contains no robot embodiment, but is widely used in embodied AI and egocentric-perception research as a robot-manipulation pretraining corpus.

present
License / format / access

Gated · Ego4D License Agreement · custom

partial
Evidence details and provenance

Official signal claims

AudioLanguageVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.