Understand activity
Temporal labels, narrations, and a task-matched evaluation split.
Research topic · Ego data
The wearer's view keeps intent, hands, objects, motion, and language close to action. It supplies human-scale supervision for embodied models, but not robot dynamics or executable controls.
Choose the supervision you need, then check coverage, access, and rights. These are editorial starting points, not a ranking or a guarantee of training suitability.
Temporal labels, narrations, and a task-matched evaluation split.
Hand and object geometry, calibration, and interaction annotations.
Paired ego–exo observations and temporal alignment.
Human demonstrations plus a method-specific embodiment bridge and robot-side evaluation.
Use and signal summaries checked against linked sources on . Access and license below come from the index and require separate confirmation.
Activity understanding, episodic memory, and forecasting.
Video · Narrations · Hand–object annotations · Audio / gaze / 3D scans in subsets
Signal source ↗Extra modalities and annotations vary by subset; check the benchmark split.
Access: Gated · License: Ego4D License Agreement
Cross-view alignment and skilled activity understanding.
Synchronized ego–exo video · Audio · Gaze · 3D
Signal source ↗Paired views support perception research; they are not robot control labels.
Access: Gated · License: Ego-Exo4D License Agreement
Human hand trajectories still need a robot-specific mapping and validation.
Access: Open · License: CC-BY-NC-ND-4.0
Geometry-aware hand–object interaction.
RGB-D · Hand pose · Object pose · Segmentation
Signal source ↗Check object categories and annotation coverage against your target task.
Access: Unknown · License: MIT
Frames show how wearer view, synchronized viewpoints, and hand pose expose different supervision.
ObserveLong-form activity in the wearer’s view
AlignFirst- and third-person skill in sync
ActHand pose for dexterous manipulationLinked dataset records remain the authority for exact rigs, calibration, sensor channels, access, and rights.
Human-reviewed index
One canonical record per human-worn first-person or synchronized ego–exo source corpus. Search and filter the source-reviewed records; access status is not a usage license.
Corpus inclusion ledger: . We check that a record identifies a human-worn first-person or synchronized ego–exo source corpus and exclude aliases and downstream repackagings.
Inclusion review does not verify every field. Access and license are reported separately; unknown means not verified, not unavailable. Open access does not grant unrestricted use.
Data-focus filters include only records with source-checked selection profiles. Other records remain searchable in All data. A missing classification is not evidence that a signal is absent. Extra modalities may exist only in subsets.
Each record links to its source for current terms and field details. Signal-source reviews have their own date; they do not refresh the corpus inclusion date.
Data focus uses source-checked profiles for 4 of 76 records; it does not search descriptions. Select All data to include unclassified records.
8 shown · 76 reviewed
Publisher not recorded · 2023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception
Inclusion ledger: 2026-08-31 · Signals not independently checked
Publisher not recorded · 2024
Aria Everyday Activities Dataset
Inclusion ledger: 2026-08-31 · Signals not independently checked
assembly-101 · 2022
Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities
Inclusion ledger: 2026-08-31 · Signals not independently checked
CaptainCook4D · 2023
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
Inclusion ledger: 2026-08-31 · Signals not independently checked
gsig · 2018
Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Inclusion ledger: 2026-08-31 · Signals not independently checked
Publisher not recorded · 2026
ChildLens: An egocentric video dataset for activity analysis in children
Inclusion ledger: 2026-08-31 · Signals not independently checked
Publisher not recorded · 2008
Guide to the Carnegie Mellon University Multimodal Activity (CMU-MMAC) Database
Inclusion ledger: 2026-08-31 · Signals not independently checked
ETH Zurich and collaborators · 2026
A dual-view egocentric interaction dataset with two synchronized head-mounted cameras, paired exocentric views, scene and object scans, actions, hand poses, gaze, and social cues.
Inclusion ledger: 2026-08-31 · Signals not independently checked
The question is no longer only whether first-person video helps. New work tests how human Ego data scales, what interaction structure must be recovered, and how much robot-side alignment still remains.
EgoScale reports 20,854 hours of action-labelled human video and a log-linear relationship between pre-training scale and validation loss, with that loss tracking downstream real-robot performance.
AoE combines a neck-mounted smartphone holder, a cross-platform capture app, and cloud–edge processing to make distributed, long-duration first-person collection a practical research pipeline.
HumanEgo lifts hand–object interaction into an embodiment-neutral representation, while EgoAVFlow models shared 3D motion for manipulation and active vision. Both move beyond copying pixels or wrist tracks alone.
EgoEngine generates robot-view observations and executable trajectories; Ego2Robot retargets actions and synthesizes robot embodiments at scale. The bridge is engineered and evaluated—not assumed.
01 · CVPR · 2022
Ego4D: Around the World in 3,000 Hours of Egocentric VideoEstablished large-scale first-person benchmarks for episodic memory, hand–object interaction, social understanding, and activity forecasting.
02 · NeurIPS · 2022
Egocentric Video-Language PretrainingBuilt EgoClip and EgoNCE to learn transferable video–language representations from narrated first-person activity.
03 · CVPR · 2022
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionAdded RGB-D sequences and geometry-aware annotations for hands, objects, poses, motion, and interaction state.
04 · CVPR · 2023
Learning Video Representations From Large Language ModelsUsed learned narrators to densify and synchronize language supervision for long-form video pre-training.
05 · ICCV · 2023
EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneMoved video–language fusion into the backbone so one pre-trained representation can support contrastive and multimodal tasks.
06 · CVPR · 2024
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesPaired ego and exo video with audio, gaze, 3D, IMU, and expert language for skill understanding and cross-view learning.
07 · CVPR · 2024
EgoExoLearn: Bridging Asynchronous Ego- and Exo-centric Procedural ActivitiesStudied how a first-person performer can follow and align with an asynchronous third-person demonstration.
Ego data supplies observation and human demonstrations. Transfer methods must still bridge viewpoint, body geometry, action space, and robot dynamics before closed-loop evaluation.
01
Human-worn RGB, depth, gaze, IMU, audio, or body sensing.
02
Hands, objects, language, temporal context, and next actions.
03
Align views, retarget motion, render robots, or distill policies.
04
Train or adapt a policy, then measure closed-loop behavior.
Official method video · silent
Ego2Robot visualizes action retargeting, robot-arm synthesis, curation, and robot-side evaluation.
Transfer evidence
Current 2026 work spans scaling, structured interaction, observation synthesis, active vision, and policy co-training. Most entries are recent preprints, so reported results are evidence to inspect—not settled consensus.
Author-reported evidence, not a cross-paper leaderboard. Tasks and evaluation settings differ.
Limit: The reported synthesized hours include morphology-specific renderings of the source behavior; they are not the same as newly recorded task hours.
Limit: The system recovers and synthesizes robot-compatible supervision; raw human RGB video is not used as an executable action stream.
Limit: The action conversion is designed for ground navigation and does not directly solve dexterous manipulation or arbitrary embodiments.
Limit: The method depends on recovering a reliable structured hand–object representation from the human video.
Limit: The shared flow representation still requires accurate 3D motion estimation and a controllable active-vision setup.
Limit: The result supports human-data pre-training at scale; it does not remove the need for robot-side alignment and closed-loop evaluation.
Retargets hand motion, removes the human arm, renders target robot embodiments, and curates the resulting trajectories at multiple quality levels.
Open paperTransforms egocentric RGB demonstrations into robot-view observation video and executable action trajectories under embodiment and feasibility constraints.
Open paperLifts hand–object interactions into an entity-level representation and trains a flow-matching policy with auxiliary interaction objectives.
Open paperUses a shared 3D flow representation to predict manipulation actions, future scene motion, and camera trajectories from egocentric demonstrations.
Open paperPre-trains on 20,854 hours of action-labelled human video, then aligns the representation with a smaller paired human–robot mid-training stage.
Open paperChoose the rig from the supervision the model needs—not from the camera name. Richer spatial and intent signals improve alignment, while adding calibration, consent, and processing cost.
Browse all indexed equipmentHead- or chest-mounted cameras; single or multi-camera rigs
Long-form video, hands, objects, speech, and activity context
Scales well, but video alone does not provide metric hand pose, depth, or robot actions.
Project Aria, Apple Vision Pro, HoloLens, and related SLAM headsets
Calibrated video, head pose, gaze, IMU, hand pose, and spatial maps
Richer geometry improves retargeting, while proprietary formats and device-specific calibration raise processing cost.
Eye tracking, IMU, EMG, pressure, touch, or synchronized body sensors
Attention, motion, muscle activity, contact, and pre-action intent cues
Adds valuable supervision, but synchronization, consent, calibration, and missing channels become first-order constraints.
6 representative systems
Start with the sensor stack, calibration path, wearer comfort, and export format. These are representative systems—not a ranking. Verify current specifications at the linked source.
Spatial research glasses
Sensor stack
RGB, tracking cameras, eye tracking, audio, and IMU
Panoramic head-mounted rig
Sensor stack
Six wide-angle RGB cameras and IMU
Head-mounted RGB-D rig
Sensor stack
Stereo depth, RGB, and IMU
Eye-tracking glasses
Sensor stack
Scene video, binocular gaze, and audio
Lightweight multi-camera headset
Sensor stack
Four synchronized RGB cameras and high-rate IMU
Neck-mounted phone kit
Sensor stack
Smartphone RGB plus app and cloud–edge pipeline
Maintain the topic
Suggest a dataset or paper, request data, or report a source error.