OpenBot
Back to Explore
Manipulation datasetOpen

CALVIN

CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

Scale
24 hours
Formats
custom
License
MIT
Published
2022-12-01

Decision summary

Best for

teleoperation

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

77/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
University of Freiburg
Evidence
secondary claim
Formats
custom
hours
24
tasks
34
bytes
656000000000
Read paper

Metadata coverage

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

depth · video · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Action / hand pose / robot state

ee_pose · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Language intent / task phase

language · has-success-labels · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Feedback / correction / failure

CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Sim-real pairing

depth · simulation · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
License / format / access

Open · MIT · custom

present
Evidence details and provenance

Official signal claims

DepthEe_poseLanguageProprioceptionTactileVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records