Language-Table
Language-Table is a suite of human-collected datasets and a multi-task continuous control benchmark for open-vocabulary visuolinguomotor learning, released with the paper 'Interactive Language: Talking to Robots in Real Time'. It contains nearly 600,000 natural-language-labeled trajectories of a UFACTORY xArm6 robot moving colored blocks on a tabletop (442,226 real-robot episodes plus simulated human-controlled and oracle episodes). Observations are third-person RGB images (captured at 640x360, resized to 320x180) plus end-effector/block state, and the arm is constrained to a 2D plane and acts via a 2D Cartesian end-effector delta setpoint.
- Scale
- 3865 hours
- Formats
- rlds
- License
- Apache-2.0
- Published
- 2022-10-12
Decision summary
scripted
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
73/100
Provisional · confidence 48 · 4/6 evaluated
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Record specifics
Dataset facts
- Source
- Robotics at Google / Google Research
- Evidence
- secondary claim
- Formats
- rlds
- episodes
- 442226
- hours
- 3865
- tasks
- 696
Metadata coverage
Loop signals
5/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
video · Language-Table is a suite of human-collected datasets and a multi-task continuous control benchmark for open-vocabulary visuolinguomotor learning, released with the paper 'Interactive Language: Talking to Robots in Real Time'. It contains nearly 600,000 natural-language-labeled trajectories of a UFACTORY xArm6 robot moving colored blocks on a tabletop (442,226 real-robot episodes plus simulated human-controlled and oracle episodes). Observations are third-person RGB images (captured at 640x360, resized to 320x180) plus end-effector/block state, and the arm is constrained to a 2D plane and acts via a 2D Cartesian end-effector delta setpoint.
presentee_pose · Language-Table is a suite of human-collected datasets and a multi-task continuous control benchmark for open-vocabulary visuolinguomotor learning, released with the paper 'Interactive Language: Talking to Robots in Real Time'. It contains nearly 600,000 natural-language-labeled trajectories of a UFACTORY xArm6 robot moving colored blocks on a tabletop (442,226 real-robot episodes plus simulated human-controlled and oracle episodes). Observations are third-person RGB images (captured at 640x360, resized to 320x180) plus end-effector/block state, and the arm is constrained to a 2D plane and acts via a 2D Cartesian end-effector delta setpoint.
partiallanguage · Language-Table is a suite of human-collected datasets and a multi-task continuous control benchmark for open-vocabulary visuolinguomotor learning, released with the paper 'Interactive Language: Talking to Robots in Real Time'. It contains nearly 600,000 natural-language-labeled trajectories of a UFACTORY xArm6 robot moving colored blocks on a tabletop (442,226 real-robot episodes plus simulated human-controlled and oracle episodes). Observations are third-person RGB images (captured at 640x360, resized to 320x180) plus end-effector/block state, and the arm is constrained to a 2D plane and acts via a 2D Cartesian end-effector delta setpoint.
presentsimulation · Language-Table is a suite of human-collected datasets and a multi-task continuous control benchmark for open-vocabulary visuolinguomotor learning, released with the paper 'Interactive Language: Talking to Robots in Real Time'. It contains nearly 600,000 natural-language-labeled trajectories of a UFACTORY xArm6 robot moving colored blocks on a tabletop (442,226 real-robot episodes plus simulated human-controlled and oracle episodes). Observations are third-person RGB images (captured at 640x360, resized to 320x180) plus end-effector/block state, and the arm is constrained to a 2D plane and acts via a 2D Cartesian end-effector delta setpoint.
presentOpen · Apache-2.0 · rlds
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
