CALVIN
CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
- Scale
- 24 hours
- Formats
- custom
- License
- MIT
- Published
- 2022-12-01
Decision summary
teleoperation
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
77/100
Provisional · confidence 48 · 4/6 evaluated
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Record specifics
Dataset facts
- Source
- University of Freiburg
- Evidence
- secondary claim
- Formats
- custom
- hours
- 24
- tasks
- 34
- bytes
- 656000000000
Metadata coverage
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
depth · video · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
presentee_pose · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
presentlanguage · has-success-labels · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
presentCALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
presentdepth · simulation · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.
presentOpen · MIT · custom
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
