OpenBot
Back to datasets
Manipulation datasetLicense requiredReadiness 70 · confidence 35

EgoSchema Dataset

Video-language model evaluation

250+hours

A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.

Best for

Video-language model evaluation

Not for / blocker

Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.

Download decision

Inspect schema and run a bounded sample audit before committing to the full release.

Policy training readinessunknown

Observation/action alignment has not been established.

not scored · confidence 15

World-model readinessunknown

World-model observation, geometry, or temporal semantics are not verified.

not scored · confidence 15

Failure and recovery readinessunknown

No verified failure/recovery annotation evidence is available yet.

not scored · confidence 15

Verified facts and provenance

Claims, metadata verification, and sample verification are shown separately.

curated source metadata
Official claim · signals
VideoLanguageLanguageTemporal Reasoning Labels
Metadata verified · schema / annotations

Unknown — no machine-readable schema facts have been captured.

Sample / pipeline verification

Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.

Declared loop signal coverage

Signals inferred from official metadata; Data pipeline verification is still pending.

5/7 categories present or partial

Observation / ego video

video · Ego4D video references · Video-language model evaluation · Underlying video access follows Ego4D licensing.

present

Action / hand pose / robot state

Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.

partial

Gaze / attention

No decision-grade evidence captured yet.

unknown

Language intent / task phase

language · temporal reasoning labels · Video-language model evaluation · A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.

present

Feedback / correction / failure

Video-language model evaluation

present

Sim-real pairing

No decision-grade evidence captured yet.

unknown

License / format / access

License required · MIT · JSON · Ego4D video references

present

Catalog decision scorecard

Access and governanceuseful

Access and license are declared by the source.

fit 75 · confidence 85

Schema and signal coverageunknown

No machine-readable schema has been verified yet.

not scored · confidence 15

Policy training readinessunknown

Observation/action alignment has not been established.

not scored · confidence 15

World-model readinessunknown

World-model observation, geometry, or temporal semantics are not verified.

not scored · confidence 15

Failure and recovery readinessunknown

No verified failure/recovery annotation evidence is available yet.

not scored · confidence 15

Download and processing readinessuseful

Scale is declared; transfer and processing estimates are not measured.

fit 65 · confidence 65

Good tasks

Video-language model evaluationLong-horizon plan checkingNarrative consistency tests for agents

Blockers and unresolved evidence

  • Gaze / attentionunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.
  • Sim-real pairingunknown
    Not enough evidence to classify this signal. Verify metadata or a bounded sample.
  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Raw dataset signals

VideoLanguageLanguageTemporal Reasoning Labels

OpenBot fit

  • Video-language model evaluation
  • Long-horizon plan checking
  • Narrative consistency tests for agents

Integration notes

  • Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
  • Underlying video access follows Ego4D licensing.

Related by signals