HABIT (Human Aware Behavior and Interaction Training dataset)
HABIT (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in human-present environments, where a co-present human shares the workspace and interacts with the…
Source notes
HABIT (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in human-present environments, where a co-present human shares the workspace and interacts with the robot in every episode.
- 10,563 episodes, 164 hours of bimanual manipulation across 60 tasks
- Organized into three human-robot interaction roles (from the HRI literature): Collaborator (jointly do one task), Coworker (separate tasks, shared space), Supervisor (human directs the robot)
- Collected with a reactive interaction protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: yielding, temporal adaptation, gesture grounding
- 5 synchronized RGB cameras (3 robot-side, 2 human-side) at 640×480, 10 Hz
- Action spaces: joint-space, Cartesian space, and grippers
- Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)
- Scale
- 164 hours
- Formats
- custom · lerobot
- License
- CC-BY-4.0
- Published
- 2026-06-30
Decision summary
teleoperation
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Release history
Source-backed release timing for this canonical dataset record.
- Release evidence
Dataset series published
Release timing is recorded from official dataset metadata.
Selection readiness
6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.
Catalog evidence · not task performance
73/100
Provisional
4/6 dimensions scored · 48% confidence
Read the evidence behind all 6 dimensions
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
- Sim-real pairing. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Dataset facts
- Source
- Config
- Evidence
- secondary claim
- Formats
- custom · lerobot
- episodes
- 10.6K
- hours
- 164
- tasks
- 60
Loop signals
4/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
video · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)
presentee_pose · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)
presentlanguage · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)
presentNo decision-grade evidence captured yet.
unknownOpen · CC-BY-4.0 · custom · lerobot
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
