OpenBot
Egocentric datasetClosed
OBRS 29Bronze

EgoSchema

A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.

Scale
250+ hours
Formats
JSON · Ego4D video references
License
MIT
Published
2023-08-17

Decision summary

Best for

Video-language model evaluation

Main blocker

Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is anchored to the cited paper publication date.

    Cited paper

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

70/100

Provisional

2/6 dimensions scored · 35% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readinessNot scoredconf. 15
World-model readinessNot scoredconf. 15
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation/action alignment has not been established.

unknownnot scored · confidence 15

World-model readiness

World-model observation, geometry, or temporal semantics are not verified.

unknownnot scored · confidence 15

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Sim-real pairing. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Dataset facts

Source
EgoSchema
Evidence
official claim
Formats
JSON · Ego4D video references
QA pairs
5,000+
hours
250+
clip length
3 min
Read paper

Loop signals

5/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Sim-real pairing

No decision-grade evidence captured yet.

unknown
View 5 more signal categories
Observation / ego video

video · Ego4D video references · Video-language model evaluation · Underlying video access follows Ego4D licensing.

present
Action / hand pose / robot state

Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.

partial
Language intent / task phase

language · temporal reasoning labels · Video-language model evaluation · A diagnostic video-language benchmark derived from Ego4D, designed to test temporal and causal reasoning over long first-person videos.

present
Feedback / correction / failure

Video-language model evaluation

present
License / format / access

Closed · MIT · JSON · Ego4D video references

present
Evidence details and provenance

Official signal claims

VideoLanguageLanguageTemporal Reasoning Labels

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated source metadata

Integration notes

  • Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
  • Underlying video access follows Ego4D licensing.