Wearable AI Dataset
Agent evaluation over streaming first-person video
A benchmark for long-form QA, conversational QA, and proactive assistant behavior over first-person wearable-camera videos.
Agent evaluation over streaming first-person video
Less robot-action focused, but useful for evaluating an agent's temporal grounding.
Inspect schema and run a bounded sample audit before committing to the full release.
Observation/action alignment has not been established.
not scored · confidence 15
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
fit 68 · confidence 50
No verified failure/recovery annotation evidence is available yet.
not scored · confidence 15
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · Agent evaluation over streaming first-person video · Conversation-grounded video retrieval · A benchmark for long-form QA, conversational QA, and proactive assistant behavior over first-person wearable-camera videos.
Action / hand pose / robot state
Less robot-action focused, but useful for evaluating an agent's temporal grounding.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · task phase
Feedback / correction / failure
Agent evaluation over streaming first-person video
Sim-real pairing
No decision-grade evidence captured yet.
License / format / access
Gated · MIT · JSONL · MP4
Catalog decision scorecard
Access and license are declared by the source.
fit 75 · confidence 85
No machine-readable schema has been verified yet.
not scored · confidence 15
Observation/action alignment has not been established.
not scored · confidence 15
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
fit 68 · confidence 50
No verified failure/recovery annotation evidence is available yet.
not scored · confidence 15
Scale is declared; transfer and processing estimates are not measured.
fit 65 · confidence 65
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Sim-real pairingunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Raw dataset signals
OpenBot fit
- Agent evaluation over streaming first-person video
- Conversation-grounded video retrieval
- Proactive assistant benchmarks
Integration notes
- Less robot-action focused, but useful for evaluating an agent's temporal grounding.
- Access requires accepting dataset terms on Hugging Face.
