EgoMM
A tri-modal egocentric dataset built from EgoLife and Ego-Exo4D sources, packaged into fixed-duration clips with optional narrations.
- Scale
- 463 hours
- Formats
- MP4 · MP3
- License
- Apache-2.0
- Published
- 2026-05-10
Decision summary
Wearable-assistant evaluation
A practical bridge between video-only egocentric data and robot-ready sensor streams.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from the official source creation date.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
70/100
Provisional
2/6 dimensions scored · 35% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation/action alignment has not been established.
World-model readiness
World-model observation, geometry, or temporal semantics are not verified.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Action / hand pose / robot state. Not enough evidence is available to classify this signal.
- Gaze / attention. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- Hugging Face · ldkong/EgoMM
- Evidence
- official claim
- Formats
- MP4 · MP3 · NPZ · JSON
- clips
- 55.5K
- hours
- 463
- narrated
- 22.1K
Loop signals
5/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
video · A practical bridge between video-only egocentric data and robot-ready sensor streams.
presentnarration · A tri-modal egocentric dataset built from EgoLife and Ego-Exo4D sources, packaged into fixed-duration clips with optional narrations.
presentWearable-assistant evaluation
presentLong-sequence reconstruction from clips
partialClosed · Apache-2.0 · MP4 · MP3
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated source metadata
Integration notes
- A practical bridge between video-only egocentric data and robot-ready sensor streams.
- Good for testing OpenBot Data's multimodal schema adapters.
