OpenBot
Manipulation datasetOpenReadiness 77/100Provisional

CALVIN

CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm.

Source notes

CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

Scale
24 hours
Formats
custom
License
MIT
Published
2022-12-01

Decision summary

Best for

teleoperation

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

How scores work

77/100

Provisional

4/6 dimensions scored · 48% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readiness85conf. 55
World-model readiness82conf. 50
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.
Engineering checks (OBRS) behind this record
Engineering checksBronze

Needs Audit

5/12checks passed

OBRS metadata and reported test evidence. A breakdown behind Selection readiness, not a separate score or an independent OpenBot certification.

Standardization & Loaders2 / 25 pt
Physical & Action Quality8 / 25 pt
Semantic & Annotation20 / 20 pt
Real-World Validation0 / 15 pt
License & Compliance10 / 15 pt
Missing evidence and readiness gaps
  • A passing real-hardware test report is required.
  • A passing ingestion/pipeline test report is required.
  • A passing privacy and provenance review is required.

Dataset facts

Source
University of Freiburg
Evidence
secondary claim
Formats
custom
hours
24
tasks
34
size
656 GB
Read paper

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

depth · video · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Action / hand pose / robot state

ee_pose · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Language intent / task phase

language · has-success-labels · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Feedback / correction / failure

CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
Sim-real pairing

depth · simulation · CALVIN (Composing Actions from Language and Vision) is an open-source PyBullet-simulated benchmark for learning long-horizon, language-conditioned continuous-control manipulation policies with a Franka Panda arm. It spans four tabletop environments (A, B, C, D), 34 manipulation tasks, and ~1000 crowd-sourced language annotations, with evaluation rollouts chaining five consecutive language-specified sub-tasks. The dataset provides ~6 hours of teleoperated 'play' data per environment (24h total) with static and gripper RGB-D cameras, tactile images, and proprioception.

present
License / format / access

Open · MIT · custom

present
Evidence details and provenance

Official signal claims

DepthEe_poseLanguageProprioceptionTactileVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.
CALVIN Dataset: License, Format & Readiness · OpenBot.ai