OpenBot
Back to Explore
Manipulation datasetOpen

HABIT (Human Aware Behavior and Interaction Training dataset)

**HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)

Scale
164 hours
Formats
custom · lerobot
License
CC-BY-4.0
Published
2026-06-30

Decision summary

Best for

teleoperation

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

73/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Feedback / correction / failure. Not enough evidence is available to classify this signal.
  • Sim-real pairing. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
Config
Evidence
secondary claim
Formats
custom · lerobot
episodes
10563
hours
164
tasks
60
Read paper

Metadata coverage

Loop signals

4/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
Sim-real pairing

No decision-grade evidence captured yet.

unknown
View 4 more signal categories
Observation / ego video

video · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)

present
Action / hand pose / robot state

ee_pose · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)

present
Language intent / task phase

language · **HABIT** (Human-Aware Behavior and Interaction Training dataset) is a large-scale robot manipulation dataset collected in **human-present** environments, where a co-present human shares the workspace and interacts with the robot in **every episode**. - **10,563 episodes**, **164 hours** of bimanual manipulation across **60 tasks** - Organized into three human-robot interaction roles (from the HRI literature): **Collaborator** (jointly do one task), **Coworker** (separate tasks, shared space), **Supervisor** (human directs the robot) - Collected with a **reactive interaction** protocol (the robot responds to the human's actual actions, not a memorized routine) and designed to elicit human-aware behaviors: **yielding, temporal adaptation, gesture grounding** - **5 synchronized RGB cameras** (3 robot-side, 2 human-side) at **640×480, 10 Hz** - Action spaces: **joint-space, Cartesian space, and grippers** - Language annotations : provide both high-level and low-level instruction per frame (both for human and robot)

present
License / format / access

Open · CC-BY-4.0 · custom · lerobot

present
Evidence details and provenance

Official signal claims

Ee_poseLanguageProprioceptionVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records