RT-2
Google DeepMind
Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.
Semantic generalization
Useful reference point for VLA direction even if the model itself is not open.
A model hub link is a declaration. Files, loadability, evaluation, and deployment are scored separately.
Model decision scorecard
Use-case scores and evidence confidence are separate; Unknown is not treated as failure.
Code, weights, and checkpoints
UnknownArtifact availability is unknown.
Loading and training reproducibility
UnknownNo verified loading configuration is available.
Training data requirements
UsefulRequired signal categories are declared; exact tensor and action interfaces still need verification.
Evaluation evidence
UsefulEvaluation focus is declared, but metrics are not independently verified.
Deployment readiness
UnknownHardware, latency, dependencies, and runtime loading are not yet verified.
Artifact facts and provenance
No metadata-verified artifact facts yet. Source links remain declarations only.
Loop signal demand
Signals this model family needs for training, evaluation, or failure mining.
Observation / ego video
observation · Language-conditioned tasks with visual observations
Language intent / task phase
language intent · task success · Language-conditioned tasks with visual observations · Instruction following under novel object combinations
Action / robot state
actions · Web-scale semantic knowledge plus robot action data · Failure cases where web knowledge does not ground to action
Future state / dynamics
Needs future-state supervision or rollout structure to validate predictive dynamics.
Feedback / correction / failure
task success · Failure cases where web knowledge does not ground to action
Sim-real / embodiment metadata
Web-scale semantic knowledge plus robot action data
Evaluation focus
- Semantic generalization
- Instruction following under novel object combinations
- Failure cases where web knowledge does not ground to action
Missing critical loop signals
Core signal demands are represented. Check quality, alignment, and access constraints.
Related catalog datasets
EPIC-KITCHENS-100
Unscripted kitchen actions from wearable cameras
Exact action dimensions, control frequency, normalization, and camera mapping require interface verification.
Ego-Exo4D
Synchronized first-person and third-person skilled activity
Exact action dimensions, control frequency, normalization, and camera mapping require interface verification.
MicroAGI01
Household manipulation with pose annotations
Exact action dimensions, control frequency, normalization, and camera mapping require interface verification.
OpenBot notes
- Useful reference point for VLA direction even if the model itself is not open.
- Highlights why OpenBot should track both language/task labels and action traces.
