OpenBot
Egocentric datasetGated
OBRS 36Bronze

Ego-Exo4D

Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria…

Source notes

Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

Scale
1,286.3 hours
Formats
custom
License
Ego-Exo4D License Agreement
Published
2023-11-30

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Ego research graph

Related research

Reviewed paper relationships connect this source record to training roles. They do not imply a general model-performance claim.

  1. Introduces dataset · CVPR · 2024

    Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Paper

    Introduces the synchronized Ego-Exo4D corpus and benchmarks.

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

69/100

Provisional

3/6 dimensions scored · 41% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readinessNot scoredconf. 15
World-model readiness68conf. 50
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation/action alignment has not been established.

unknownnot scored · confidence 15

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

    Next checks

    • Read the dataset manifest and feature schema.
    • Run a bounded sample audit before assigning Strong readiness.

    Dataset facts

    Source
    Meta FAIR (Project Aria) and 15+ university partners
    Evidence
    secondary claim
    Formats
    custom
    episodes
    5,035
    hours
    1,286.3
    tasks
    4
    Read paper

    Loop signals

    7/7 present or partial

    Observation / ego video

    video · Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    present
    Action / hand pose / robot state

    Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    partial
    View 5 more signal categories
    Gaze / attention

    Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    present
    Language intent / task phase

    language · Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    present
    Feedback / correction / failure

    Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    present
    Sim-real pairing

    Ego-Exo4D is a large-scale multimodal, multiview video dataset of skilled human activities (cooking, music, soccer, health, basketball, dance, bike repair, rock climbing) captured simultaneously from first-person Aria glasses and third-person GoPro cameras. It contains 1,286.3 hours of video from 740 camera wearers across 13 cities and 123 scene contexts, with multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions including expert commentary. It ships four benchmark tasks: keystep/fine-grained activity recognition, proficiency estimation, ego-exo relation/cross-view translation, and 3D hand/body pose estimation.

    present
    License / format / access

    Gated · Ego-Exo4D License Agreement · custom

    partial
    Evidence details and provenance

    Official signal claims

    AudioLanguagePoint_cloudVideo

    Schema and annotations

    No machine-readable schema facts are captured.

    Sample verification

    Pending. Metadata does not prove sample coverage, alignment, or file integrity.

    curated official source

    Integration notes

    • Metadata requires review against the official source before publication.