RT-2
Google DeepMind
Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.
- Code
- Not listed
- Weights
- Not listed
- Checkpoint
- Not listed
- License
- Unverified
Decision summary
Semantic generalization
Hardware, latency, dependencies, and checkpoint loading are not pipeline-tested.
Verify repository and artifact licenses.
Release history
Source-backed release timing for this canonical model record.
- Paper
Model series introduced
Release timing is anchored to the cited paper publication date.
Evidence profile
A visual read of adoption evidence. Scores describe catalog evidence readiness, not task performance.
Published evaluation
1/6
dimensions scored
Unknown evidence remains visible and is never treated as a zero.
Why these scores
The strongest decision reasons behind the evidence profile.
Artifact availability
UnknownCode and weights are not verified.
Training and loading reproducibility
UnknownNo verified loading configuration is available.
Evaluation evidence
UnknownNo structured evaluation evidence has been verified.
Data requirements
Declared loop-data needs, missing evidence, and the next checks that matter.
Observation / ego video
observation
Language intent / task phase
language intent
Action / robot state
actions
Feedback / correction / failure
task success
Critical gaps
Core categories are represented. Interface alignment and data quality still require verification.
Linked datasets
Signal-level links only. Verify runtime interfaces before use.
More evidenceShowHide
Additional data signals
Sim-real / embodiment metadata
Descriptive model record
Additional decision dimensions
License or access terms are not verified.
— · c20Required signal categories are structured; exact tensor and action interfaces still need verification.
76 · c70Hardware, latency, dependencies, and checkpoint loading are not pipeline-tested.
— · c10Artifact facts
- release_year
- official_claim
- 2023
- curated official source year
- model.required_signals
- official_claim
- observation · language intent · actions · object labels · task success
- Official model documentation and curated signal mapping
- release_timing
- metadata_verified
- 2023-07-28T21:18:02.000Z
- arXiv 2307.15818 published
- discovery.source
- secondary_claim
- name: worldbench/awesome-embodied-data-pyramid · trust: discovery_only · revision: 0568599d0619f20946e8089f6a941d0d9e30b690
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- discovery.raw_fields
- secondary_claim
- Time: 2023.7 · Method: RT-2 · Institution: DeepMind · Project: [](https://robotics-transformer2.github.io/) · Model: VLA · Data: ![Real][data-real] ![General][data-general]
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- model.reported_release_time
- secondary_claim
- 2023.7
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- model.type
- secondary_claim
- VLA
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- model.training_data_layers
- secondary_claim
- real_robot · general
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- model.institution
- secondary_claim
- DeepMind
- #table-embodied-foundation-models-vla-wam, table 17, row 2
- catalog.curation
- secondary_claim
- tier: editorial_focus · collection: WorldBench Awesome Embodied Data Pyramid · policy: human_curated_priority
- #table-embodied-foundation-models-vla-wam, table 17, row 2
Additional references
OpenBot notes
- Useful reference point for VLA direction even if the model itself is not open.
- Highlights why OpenBot should track both language/task labels and action traces.
