OpenBot

Research topic · Ego data

Ego data for embodied learning.

The wearer's view keeps intent, hands, objects, motion, and language close to action. It supplies human-scale supervision for embodied models, but not robot dynamics or executable controls.

76 reviewed records4 current 2026 signals6 representative devices6 2026 transfer studies

Start with your task

Choose the supervision you need, then check coverage, access, and rights. These are editorial starting points, not a ranking or a guarantee of training suitability.

Understand activity

Temporal labels, narrations, and a task-matched evaluation split.

Compare starting datasets

Use and signal summaries checked against linked sources on . Access and license below come from the index and require separate confirmation.

Activity understanding, episodic memory, and forecasting.

Video · Narrations · Hand–object annotations · Audio / gaze / 3D scans in subsets

Signal source ↗

Extra modalities and annotations vary by subset; check the benchmark split.

Access: Gated · License: Ego4D License Agreement

Cross-view alignment and skilled activity understanding.

Synchronized ego–exo video · Audio · Gaze · 3D

Signal source ↗

Paired views support perception research; they are not robot control labels.

Access: Gated · License: Ego-Exo4D License Agreement

Dexterous human demonstration and hand-motion research.

Video · 3D hand tracking

Signal source ↗

Human hand trajectories still need a robot-specific mapping and validation.

Access: Open · License: CC-BY-NC-ND-4.0

Geometry-aware hand–object interaction.

RGB-D · Hand pose · Object pose · Segmentation

Signal source ↗

Check object categories and annotation coverage against your target task.

Access: Unknown · License: MIT

See the supervision

Frames show how wearer view, synchronized viewpoints, and hand pose expose different supervision.

Linked dataset records remain the authority for exact rigs, calibration, sensor channels, access, and rights.

Human-reviewed index

Source datasets

One canonical record per human-worn first-person or synchronized ego–exo source corpus. Search and filter the source-reviewed records; access status is not a usage license.

What has been reviewed?

Corpus inclusion ledger: . We check that a record identifies a human-worn first-person or synchronized ego–exo source corpus and exclude aliases and downstream repackagings.

Inclusion review does not verify every field. Access and license are reported separately; unknown means not verified, not unavailable. Open access does not grant unrestricted use.

Data-focus filters include only records with source-checked selection profiles. Other records remain searchable in All data. A missing classification is not evidence that a signal is absent. Extra modalities may exist only in subsets.

Each record links to its source for current terms and field details. Signal-source reviews have their own date; they do not refresh the corpus inclusion date.

Data focus uses source-checked profiles for 4 of 76 records; it does not search descriptions. Select All data to include unclassified records.

8 shown · 76 reviewed

  1. ADTAccess not verified

    Publisher not recorded · 2023

    Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  2. AEAAccess not verified

    Publisher not recorded · 2024

    Aria Everyday Activities Dataset

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  3. Assembly101Access not verified

    assembly-101 · 2022

    Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  4. CaptainCook4DAccess not verified

    CaptainCook4D · 2023

    CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  5. Charades-EgoAccess not verified

    gsig · 2018

    Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  6. ChildLensAccess not verified

    Publisher not recorded · 2026

    ChildLens: An egocentric video dataset for activity analysis in children

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  7. CMU-MMACAccess not verified

    Publisher not recorded · 2008

    Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) Database

    Catalog-reported signals
    Not verified
    Reported license
    Not verified — check source terms

    Inclusion ledger: 2026-08-31 · Signals not independently checked

  8. CoMindOpen

    ETH Zurich and collaborators · 2026

    A dual-view egocentric interaction dataset with two synchronized head-mounted cameras, paired exocentric views, scene and object scans, actions, hand poses, gaze, and social cues.

    Catalog-reported signals
    Video / RGB · Depth / 3D · Pose / calibration · Objects / hands · Audio
    Reported license
    CC-BY-4.0

    Inclusion ledger: 2026-08-31 · Signals not independently checked

Why Ego matters in 2026

The question is no longer only whether first-person video helps. New work tests how human Ego data scales, what interaction structure must be recovered, and how much robot-side alignment still remains.

  1. Human data now shows a scaling signal

    EgoScale reports 20,854 hours of action-labelled human video and a log-linear relationship between pre-training scale and validation loss, with that loss tracking downstream real-robot performance.

  2. Collection can move beyond the lab

    AoE combines a neck-mounted smartphone holder, a cross-platform capture app, and cloud–edge processing to make distributed, long-duration first-person collection a practical research pipeline.

    EvidenceAoE · 2026
  3. The transferable unit is structured interaction

    HumanEgo lifts hand–object interaction into an embodiment-neutral representation, while EgoAVFlow models shared 3D motion for manipulation and active vision. Both move beyond copying pixels or wrist tracks alone.

  4. The embodiment gap is now an explicit stage

    EgoEngine generates robot-view observations and executable trajectories; Ego2Robot retargets actions and synthesizes robot embodiments at scale. The bridge is engineered and evaluated—not assumed.

Evidence baseFoundational Ego papers7 papers
  1. 01 · CVPR · 2022

    Ego4D: Around the World in 3,000 Hours of Egocentric Video

    Established large-scale first-person benchmarks for episodic memory, hand–object interaction, social understanding, and activity forecasting.

  2. 02 · NeurIPS · 2022

    Egocentric Video-Language Pretraining

    Built EgoClip and EgoNCE to learn transferable video–language representations from narrated first-person activity.

  3. 03 · CVPR · 2022

    HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

    Added RGB-D sequences and geometry-aware annotations for hands, objects, poses, motion, and interaction state.

  4. 04 · CVPR · 2023

    Learning Video Representations From Large Language Models

    Used learned narrators to densify and synchronize language supervision for long-form video pre-training.

  5. 05 · ICCV · 2023

    EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

    Moved video–language fusion into the backbone so one pre-trained representation can support contrastive and multimodal tasks.

  6. 06 · CVPR · 2024

    Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Paired ego and exo video with audio, gaze, 3D, IMU, and expert language for skill understanding and cross-view learning.

  7. 07 · CVPR · 2024

    EgoExoLearn: Bridging Asynchronous Ego- and Exo-centric Procedural Activities

    Studied how a first-person performer can follow and align with an asynchronous third-person demonstration.

From human view to robot learning

Ego data supplies observation and human demonstrations. Transfer methods must still bridge viewpoint, body geometry, action space, and robot dynamics before closed-loop evaluation.

  1. 01

    Capture signals

    Human-worn RGB, depth, gaze, IMU, audio, or body sensing.

  2. 02

    Learn state and intent

    Hands, objects, language, temporal context, and next actions.

  3. 03

    Bridge embodiment

    Align views, retarget motion, render robots, or distill policies.

  4. 04

    Validate on a robot

    Train or adapt a policy, then measure closed-loop behavior.

Official method video · silent

From human demonstrations to robot training data

Ego2Robot visualizes action retargeting, robot-arm synthesis, curation, and robot-side evaluation.

Project page

Transfer evidence

2026 Ego-to-Robot studies

6 studies · 2026

Current 2026 work spans scaling, structured interaction, observation synthesis, active vision, and policy co-training. Most entries are recent preprints, so reported results are evidence to inspect—not settled consensus.

Compare transfer requirements

Author-reported evidence, not a cross-paper leaderboard. Tasks and evaluation settings differ.

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Input
EgoDex; EgoVerse; ViTRA; ANT
Method
Robot-data synthesis
Robot demonstrations
Joint pre-training with robot data
Validation
VLA pre-training across morphology and scene perturbations, followed by real-robot deployment.

Limit: The reported synthesized hours include morphology-specific renderings of the source behavior; they are not the same as newly recorded task hours.

EgoEngine: From Egocentric Human Videos to Robot Learning
Input
Egocentric human RGB demonstrations
Method
Robot-data synthesis
Robot demonstrations
No target-task real-robot demonstrations reported
Validation
Policies are evaluated in simulation and on real robots, including zero-shot visuomotor dexterous manipulation.

Limit: The system recovers and synthesizes robot-compatible supervision; raw human RGB video is not used as an executable action stream.

Co-training with Ego-centric Video and Demonstration for Robot Navigation Task
Input
Egocentric human walking videos; Robot navigation demonstrations
Method
Policy co-training
Robot demonstrations
Human-derived data plus robot demonstrations
Validation
Evaluated on a fruit-search navigation task with a mobile robot.

Limit: The action conversion is designed for ground navigation and does not directly solve dexterous manipulation or arbitrary embodiments.

HumanEgo: Learning Robot Manipulation from Egocentric Human Videos
Input
Purpose-recorded egocentric human demonstrations
Method
Structured policy learning
Robot demonstrations
No robot demonstrations reported
Validation
Reports zero-shot transfer across real robot hardware using short human demonstration collections.

Limit: The method depends on recovering a reliable structured hand–object representation from the human video.

EgoAVFlow: Learning Unified 3D Flow for Robot Manipulation and Active Vision
Input
Egocentric human manipulation videos
Method
Active-view policy
Robot demonstrations
No robot demonstrations reported
Validation
Real-world experiments evaluate manipulation under viewpoint changes and active camera motion.

Limit: The shared flow representation still requires accurate 3D motion estimation and a controllable active-vision setup.

EgoScale: Scaling Egocentric Human Data for Robot Learning
Input
20,854 hours of action-labelled egocentric human video; Aligned human–robot demonstrations
Method
Human-data scaling
Robot demonstrations
Small aligned human–robot set plus downstream robot data
Validation
Evaluated on real-robot manipulation with a 22-DoF robotic hand across multiple tasks.

Limit: The result supports human-data pre-training at scale; it does not remove the need for robot-side alignment and closed-loop evaluation.

Robot-data synthesis · arXiv · 2026Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Retargets hand motion, removes the human arm, renders target robot embodiments, and curates the resulting trajectories at multiple quality levels.

Open paper
Robot-data synthesis · arXiv · 2026EgoEngine: From Egocentric Human Videos to Robot Learning

Transforms egocentric RGB demonstrations into robot-view observation video and executable action trajectories under embodiment and feasibility constraints.

Open paper
Policy co-training · arXiv · 2026Co-training with Ego-centric Video and Demonstration for Robot Navigation Task

Converts human walking video into ground-mobile-robot action representations and co-trains a vision–language–action model with robot demonstrations.

Open paper
Structured policy learning · arXiv · 2026HumanEgo: Learning Robot Manipulation from Egocentric Human Videos

Lifts hand–object interactions into an entity-level representation and trains a flow-matching policy with auxiliary interaction objectives.

Open paper
Active-view policy · arXiv · 2026EgoAVFlow: Learning Unified 3D Flow for Robot Manipulation and Active Vision

Uses a shared 3D flow representation to predict manipulation actions, future scene motion, and camera trajectories from egocentric demonstrations.

Open paper
Human-data scaling · arXiv · 2026EgoScale: Scaling Egocentric Human Data for Robot Learning

Pre-trains on 20,854 hours of action-labelled human video, then aligns the representation with a smaller paired human–robot mid-training stage.

Open paper

What to capture

Choose the rig from the supervision the model needs—not from the camera name. Richer spatial and intent signals improve alignment, while adding calibration, consent, and processing cost.

Browse all indexed equipment

Wearable RGB

Head- or chest-mounted cameras; single or multi-camera rigs

Long-form video, hands, objects, speech, and activity context

Scales well, but video alone does not provide metric hand pose, depth, or robot actions.

Spatial headset

Project Aria, Apple Vision Pro, HoloLens, and related SLAM headsets

Calibrated video, head pose, gaze, IMU, hand pose, and spatial maps

Richer geometry improves retargeting, while proprietary formats and device-specific calibration raise processing cost.

RGB-D interaction rig

Synchronized color and depth cameras, including head-mounted RGB-D

Depth, 3D hand and object pose, segmentation, and interaction geometry

Better for contact and geometry, but heavier rigs constrain natural motion and collection scale.

Intent and body sensing

Eye tracking, IMU, EMG, pressure, touch, or synchronized body sensors

Attention, motion, muscle activity, contact, and pre-action intent cues

Adds valuable supervision, but synchronization, consent, calibration, and missing channels become first-order constraints.

Capture equipment index

6 representative systems

Start with the sensor stack, calibration path, wearer comfort, and export format. These are representative systems—not a ranking. Verify current specifications at the linked source.

Project Aria Gen 2

Spatial research glasses

Sensor stack

RGB, tracking cameras, eye tracking, audio, and IMU

Intended use

Calibrated multimodal capture and spatial reconstruction

Community index

DAS Ego

Panoramic head-mounted rig

Sensor stack

Six wide-angle RGB cameras and IMU

Intended use

Wide-field hand, object, and environment coverage

Official product

Orbbec EGO RGB-D

Head-mounted RGB-D rig

Sensor stack

Stereo depth, RGB, and IMU

Intended use

Metric geometry, contact, and hand–object reconstruction

Manufacturer site

Tobii Pro Glasses 3

Eye-tracking glasses

Sensor stack

Scene video, binocular gaze, and audio

Intended use

Attention, intent, and task-analysis studies

Official product

LANYVE-OMNI 4P

Lightweight multi-camera headset

Sensor stack

Four synchronized RGB cameras and high-rate IMU

Intended use

Wide-view capture where wearer comfort matters

Manufacturer site

AoE smartphone rig

Neck-mounted phone kit

Sensor stack

Smartphone RGB plus app and cloud–edge pipeline

Intended use

Low-barrier, distributed, long-duration collection

Research paper

Maintain the topic

Suggest a dataset or paper, request data, or report a source error.

contact@openbot.ai