Xperience-10M
A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
- Scale
- 10K hours
- Formats
- Hugging Face dataset · multimodal episode files
- License
- Custom terms · approved non-commercial use
- Published
- 2026-03-11
Decision summary
World model pretraining
Very large and controlled-access; the practical OpenBot path is metadata indexing plus targeted subset pulls.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from the official source creation date.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
77/100
Provisional
4/6 dimensions scored · 48% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- Hugging Face · ropedia-ai/xperience-10m
- Evidence
- official claim
- Formats
- Hugging Face dataset · multimodal episode files
- experiences
- 10M
- hours
- 10K
- RGB frames
- 2.88B
Loop signals
5/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
video · depth · camera pose · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
presentcamera pose · hand pose · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
partiallanguage · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
presentdepth · camera pose · Real-to-sim and sim-to-real data alignment · A large egocentric multimodal dataset with synchronized video streams, audio, depth, poses, mocap, IMU, and hierarchical language annotations.
presentGated · Custom terms · approved non-commercial use · Hugging Face dataset · multimodal episode files
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated source metadata
Integration notes
- Very large and controlled-access; the practical OpenBot path is metadata indexing plus targeted subset pulls.
- Useful as a reference for the signals OpenBot Data should preserve.
