EO-Data1.5M
A large multimodal corpus combining vision, text, embodied reasoning, and action supervision for EO-1 pretraining.
- Scale
- 1.5M samples
- Formats
- Hugging Face dataset
- License
- Not declared
- Published
- 2025-08-28
Decision summary
Embodied reasoning
Separate EO-Data1.5M training samples from EO-Bench evaluation data.
Verify repository and artifact licenses.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
78/100
Provisional
3/6 dimensions scored · 37% confidence
Read the evidence behind all 6 dimensions
Access and governance
License or access terms are not verified.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
A bounded sample/subset is declared for initial validation.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
- Sim-real pairing. Not enough evidence is available to classify this signal.
Next checks
- Verify repository and artifact licenses.
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- IPEC-COMMUNITY / EO-Robotics
- Evidence
- official claim
- Formats
- Hugging Face dataset
- samples
- 1.5M
- role
- Pretraining
- modalities
- Vision + text + action
Loop signals
5/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
video · images
presentactions · future state · A large multimodal corpus combining vision, text, embodied reasoning, and action supervision for EO-1 pretraining.
presentlanguage · task phase
presentSeparate EO-Data1.5M training samples from EO-Bench evaluation data.
presentOpen · Hugging Face dataset
partialEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated source metadata
Integration notes
- Separate EO-Data1.5M training samples from EO-Bench evaluation data.
