OpenBot
Models & papers38 models · 5 families · fallback

Embodied foundation models need loop data.

A curated map of model families, links, and the loop signals each one needs.

Reading frame
Data requirements, not hype
Families
VLA · World Model · WAM · Policy
Core signals
observation · intent · action · feedback
Catalog link
dataset coverage becomes model readiness
Output
what to train, evaluate, or collect next
Signal coverage

Papers point back to missing data.

Each model is a demand signal for catalog fields, curation, evaluation cases, and replay.

38 model refs
Observation
33 model refs
Actions
27 model refs
Language intent
24 model refs
Robot state
10 model refs
Task phase
9 model refs
Future state
7 model refs
Feedback/failure
7 model refs
Task success
6 model refs
Depth
6 model refs
Sim-real
5 model refs
Camera pose
3 model refs
Object labels
Model index

Search models by family, openness, and loop-data demand.

A practical map for finding which datasets support training, world modeling, WAM-style action, or failure mining.

Filters

Showing 38 of 38 models

WAMResearch2026

DreamZero / World Action Models

Research

Detail

World-action model framing where future visual states and actions are predicted in an aligned policy model.

Recommended for

Zero-shot policy behavior

Current blocker

This is the strongest reason to keep Catalog focused on loop traces instead of flat dataset metadata.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Observation, intent, action, and future-state supervision
  • Robot demonstrations with paired video/action sequences
  • Failure and correction signals for validating imagined futures
Dataset signals
ObservationActionsFuture stateLanguage intentFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026model hub

GR00T 1.7

NVIDIA

Detail

NVIDIA's current open humanoid VLA release within an end-to-end Isaac workflow for teleoperation, training, evaluation, and deployment.

Recommended for

Humanoid cross-embodiment deployment

Current blocker

Track as the latest release in the GR00T series, not as a replacement for historical N1 results.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Human demonstrations
  • Simulation trajectories
  • Whole-body robot state
Dataset signals
ObservationLanguage intentActionsRobot stateSim-realTask phase
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAClosed ref2026

Helix 02

Figure

Detail

A hierarchical whole-body VLA for continuous humanoid locomotion, manipulation, balance, tactile feedback, and long-horizon autonomy.

Recommended for

Whole-body loco-manipulation

Current blocker

Highlights why Catalog needs humanoid, tactile, and whole-body signal fields.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Whole-body human motion
  • Humanoid proprioception
  • Palm and head cameras
Dataset signals
ObservationTactileRobot stateActionsFeedback/failureTask phase
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
PerceptionOpen2026model hub

LingBot-Depth

Robbyant

Detail

A masked depth model for refining incomplete sensor depth into metric geometry for reconstruction, tracking, and manipulation.

Recommended for

Metric depth accuracy

Current blocker

The recommended v0.5 release supersedes the earlier v0.1 checkpoint.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Calibrated RGB-D pairs
  • Raw and refined metric depth
  • Camera intrinsics
Dataset signals
ObservationDepthCamera CalibrationPoint CloudSim-real
Model hubDataset
curated source· verification pending· code declared· weights declared
Evaluate in Bench
PerceptionOpen2026model hub

LingBot-Map

Robbyant

Detail

A feed-forward 3D foundation model for streaming scene reconstruction, camera-pose estimation, and long-sequence point-cloud generation.

Recommended for

Streaming reconstruction

Current blocker

Long, balanced, and Stage-1 checkpoints belong to one model series.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Streaming image sequences
  • Camera trajectories
  • Metric geometry
Dataset signals
ObservationCamera poseDepthPoint CloudTrajectory
Model hubDataset
curated source· verification pending· code declared· weights declared
Evaluate in Bench
WAMOpen2026model hub

LingBot-VA

Robbyant

Detail

An autoregressive video-action world-model policy that interleaves future video-latent prediction with robot action generation.

Recommended for

Video-conditioned action prediction

Current blocker

Available through the LeRobot policy interface with public base and post-trained checkpoints.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Aligned video and action sequences
  • Future-state supervision
  • Task text and action normalization
Dataset signals
ObservationFuture stateActionsRobot stateLanguage intent
DocsModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
PerceptionOpen2026model hub

LingBot-Vision

Robbyant

Detail

A family of vision foundation encoders pretrained for dense spatial perception and downstream embodied understanding.

Recommended for

Dense spatial perception

Current blocker

Small, Base, Large, and Giant are checkpoints in one vision model series.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large-scale visual pretraining data
  • Dense spatial labels
  • Depth and geometry benchmarks
Dataset signals
ObservationDepthObjectsPoseCamera Calibration
Model hubModel hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAOpen2026model hub

LingBot-VLA

Robbyant

Detail

A 4B pragmatic VLA foundation model pretrained on roughly 20,000 hours of real-world data from nine dual-arm robot configurations.

Recommended for

Cross-embodiment transfer

Current blocker

Released in depth-free and depth-distilled configurations.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large-scale dual-arm demonstrations
  • Cross-embodiment action normalization
  • Optional depth supervision
Dataset signals
ObservationDepthLanguage intentActionsRobot state
Model hubModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
VLAOpen2026model hub

LingBot-VLA 2.0

Robbyant

Detail

A whole-body VLA release focused on broader cross-embodiment transfer, mobile manipulation, and predictive dynamics supervision.

Recommended for

Whole-body action generation

Current blocker

Track as a release in the LingBot-VLA series rather than an unrelated model.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Whole-body action traces
  • Multiple robot configurations
  • Human video paired with robot data
Dataset signals
ObservationFuture stateLanguage intentActionsRobot stateTask phase
Model hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
World ModelOpen2026model hub

LingBot-World

Robbyant

Detail

An open interactive world simulator derived from video generation, with camera- and action-conditioned variants and long-horizon generation.

Recommended for

Long-horizon consistency

Current blocker

Base (Cam), Base (Act), and Fast are releases/checkpoints of one model series.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Action- or camera-conditioned video
  • Long temporal sequences
  • Consistent scene dynamics
Dataset signals
ObservationFuture stateActionsCamera poseLanguage intent
Model hubModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
World ModelResearch2026

Looped World Models

Facemind / research

Detail

World-model architecture using looped transformer refinement over latent states for iterative environment understanding.

Recommended for

Data efficiency

Current blocker

Important trend signal for why OpenBot should model feedback and correction, not only observations.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Temporal egocentric and human-centric observations
  • Dense signals that expose feedback and state correction
  • Benchmarks that measure data efficiency and rollout stability
Dataset signals
ObservationFeedbackFuture stateGaze/attentionHuman demonstration
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
WAMResearch2026

OA-WAM

Research

Detail

Object-addressable world-action model that decomposes scenes into robot and object slots while jointly predicting future world state and actions.

Recommended for

Object identity under scene shifts

Current blocker

Good example of why object-level and task-phase signals matter for OpenBot dataset detail pages.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Object-centric interaction traces
  • Language instructions tied to particular objects
  • Visual, proprioceptive, and action tokens over time
Dataset signals
ObservationObject labelsActionsRobot stateLanguage intent
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAOpen2026

StarVLA

StarVLA community

Detail

A modular, MIT-licensed VLA development stack with multiple VLM backbones, model scales, and dataset integrations.

Recommended for

Architecture and backbone comparison

Current blocker

Distinguish the StarVLA framework from individual StarVLA-alpha model releases.

Code, weights, and checkpoints
58confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • LeRobot-compatible demonstrations
  • Multi-view observations
  • Language instructions
Dataset signals
ObservationLanguage intentActionsRobot stateTask phase
curated source· verification pending· code declared· weights unknown
Evaluate in Bench
VLAOpen2026model hub

WALL-OSS 0.5

X Square Robot

Detail

An open 4B VLA pretrained across more than 20 embodiments, with physical-hardware evaluation of pretrained robotic capability.

Recommended for

Zero-shot physical capability

Current blocker

Officially integrated into LeRobot; self-reported results should remain separate from independent reproduction.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large cross-embodiment trajectory mixtures
  • Grounded multimodal data
  • Task decomposition and continuous actions
Dataset signals
ObservationLanguage intentTask phaseActionsRobot state
Model hubDocs
curated source· verification pending· code declared· weights declared
Evaluate in Bench
WAMOpen2025model hub

EO-1

IPEC / EO-Robotics

Detail

A unified embodied model using interleaved vision-text-action pretraining for perception, reasoning, planning, and continuous robot control.

Recommended for

Unified reasoning and action

Current blocker

Paired with EO-Data1.5M and EO-Bench; keep training and benchmark relations explicit.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Interleaved vision-text-action samples
  • Embodied reasoning traces
  • Continuous robot actions
Dataset signals
ObservationLanguage intentTask phaseActionsFuture state
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAOpen2025model hub

Galaxea G0

OpenGalaxea

Detail

A dual-system VLM plus VLA model for planning and fine-grained control on long-horizon mobile-manipulation tasks.

Recommended for

Long-horizon mobile manipulation

Current blocker

Directly paired with the Galaxea Open-World Dataset.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Subtask-level language
  • Mobile dual-arm trajectories
  • Long-horizon task structure
Dataset signals
ObservationLanguage intentTask phaseActionsRobot stateOutcome
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAClosed ref2025

Gemini Robotics

Google DeepMind

Detail

Embodied reasoning and action model family that brings Gemini capabilities into robotics through vision, language, and physical action.

Recommended for

Physical reasoning and instruction following

Current blocker

A closed reference point, but useful for tracking the VLA/embodied foundation model direction.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Vision-language interaction data connected to executable robot skills
  • Generalization tests across objects, scenes, and task instructions
  • Safety, affordance, and failure feedback for physical-world deployment
Dataset signals
ObservationLanguage intentActionsObject labelsFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAClosed ref2025

Gemini Robotics 1.5

Google DeepMind

Detail

An agentic robotics model combining advanced multimodal reasoning with vision-language-action control for multi-step physical tasks.

Recommended for

Multi-step reasoning

Current blocker

Keep availability distinct from Gemini Robotics On-Device and ER releases.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Multi-step instructions
  • Visual observations
  • Action traces
Dataset signals
ObservationLanguage intentTask phaseActionsFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAOpen2025model hub

GR00T N1

NVIDIA

Detail

Open foundation model for generalist humanoid robots, focused on whole-body and manipulation behavior from multimodal robot data.

Recommended for

Generalist humanoid task transfer

Current blocker

Pushes OpenBot tags toward embodiment-aware metadata rather than generic video/action labels.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Humanoid or whole-body demonstrations with proprioception, actions, and visual context
  • Task-level language or intent metadata
  • Sim-real and embodiment metadata for evaluating transfer to physical robots
Dataset signals
ObservationLanguage intentActionsRobot stateSim-real
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
World ModelOpen2025model hub

NVIDIA Cosmos

NVIDIA

Detail

World foundation model platform for physical AI, built around predictive world modeling and data processing workflows.

Recommended for

Future-state prediction quality

Current blocker

Useful anchor for the world-model side of OpenBot's Catalog and Synth story.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large-scale video and interaction data
  • Geometry, depth, calibration, or simulation context
  • Evaluation traces that connect prediction quality to downstream policy improvement
Dataset signals
ObservationDepthCamera poseSim-realFuture state
Model hubModel hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAOpen2025

OpenVLA-OFT

OpenVLA collaboration

Detail

An optimized OpenVLA fine-tuning recipe with continuous action chunks, multi-image input, and faster high-frequency control.

Recommended for

Fine-tuning speed

Current blocker

Treat as an OpenVLA release/recipe in the future series schema, despite the current flat static index.

Code, weights, and checkpoints
58confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Task-specific robot demonstrations
  • Multiple camera views
  • Continuous action chunks
Dataset signals
ObservationLanguage intentActionsRobot stateCamera pose
curated source· verification pending· code declared· weights unknown
Evaluate in Bench
VLAResearch2025

pi0.5

Physical Intelligence

Detail

A VLA model designed for open-world generalization by combining heterogeneous robot data, semantic subtask prediction, and high-level knowledge transfer.

Recommended for

Open-world generalization

Current blocker

A useful anchor for Catalog fields around task phase and heterogeneous data mixtures.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Heterogeneous robot demonstrations
  • High-level subtask labels
  • Language instructions
Dataset signals
ObservationLanguage intentTask phaseActionsRobot state
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAOpen2025model hub

SmolVLA

Hugging Face LeRobot

Detail

Compact open vision-language-action policy designed for practical robot fine-tuning and deployment through the LeRobot ecosystem.

Recommended for

Data-efficient fine-tuning

Current blocker

Strong fit for OpenBot's model-readiness view because public weights make the training path concrete.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • LeRobot-style trajectories with observations, actions, and task text
  • Small but clean demonstrations that preserve episode boundaries and action timing
  • Evaluation splits that measure fine-tuning data efficiency
Dataset signals
ObservationLanguage intentActionsRobot stateTask success
PostModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
VLAOpen2025model hub

SpatialVLA

Shanghai AI Laboratory / collaborators

Detail

A spatially enhanced 4B VLA pretrained on 1.1 million real-robot episodes with explicit geometry-aware representations.

Recommended for

Spatial grounding

Current blocker

Official training references Open X-Embodiment and RH20T.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large-scale real robot episodes
  • Depth and spatial supervision
  • RLDS action/state alignment
Dataset signals
ObservationDepthLanguage intentActionsRobot state
Model hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
VLAOpen2025model hub

X-VLA

Tsinghua AIR / collaborators

Detail

A 0.9B soft-prompted VLA that adapts a shared backbone across heterogeneous embodiments and action spaces.

Recommended for

Cross-embodiment transfer

Current blocker

Natively integrated into LeRobot with public base and benchmark checkpoints.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Cross-embodiment datasets
  • Explicit domain and action-mode metadata
  • Robot-specific post-training demonstrations
Dataset signals
ObservationLanguage intentActionsRobot stateEmbodiment Metadata
DocsModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
World ModelClosed ref2024

Genie 2

Google DeepMind

Detail

Large-scale foundation world model for generating action-controllable interactive environments from visual prompts.

Recommended for

Controllable world generation

Current blocker

Not a robot policy, but a useful trend marker for the world-model side of embodied AI.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Action-controllable video or environment interaction traces
  • Consistent observations over time with enough geometry and dynamics
  • Benchmarks that test controllability, persistence, and physical plausibility
Dataset signals
ObservationActionsFuture stateSim-realFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
PolicyOpen2024model hub

Octo

Octo Model Team

Detail

Open-source generalist robot policy pretrained on Open X-Embodiment trajectories and designed for fine-tuning to new robots and tasks.

Recommended for

Fine-tuning to new observation spaces

Current blocker

Good bridge between pure policy learning and broader VLA/WAM framing.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Robot trajectories with observations and actions
  • Flexible task definitions such as language or goal images
  • Sensor/action-space metadata for adaptation
Dataset signals
ObservationActionsRobot stateLanguage intentGoal image
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
VLAOpen2024model hub

OpenVLA

Stanford / UC Berkeley / Toyota Research Institute / collaborators

Detail

Open-source vision-language-action model for generalist robotic manipulation, trained on diverse real-world robot demonstrations.

Recommended for

Task success across objects and scenes

Current blocker

A strong anchor for mapping dataset readiness to VLA fine-tuning needs.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Robot demonstrations with image observations and action traces
  • Language instructions or task labels aligned to episodes
  • Embodiment metadata for fine-tuning and evaluation splits
Dataset signals
ObservationLanguage intentActionsRobot stateTask success
Model hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
VLAOpen2024model hub

pi0 / OpenPI

Physical Intelligence

Detail

Generalist robot policy family and open-source robotics model package from Physical Intelligence.

Recommended for

Generalist policy transfer

Current blocker

Useful for framing OpenBot as a data-readiness layer around emerging generalist policies.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Broad robot demonstrations with language goals
  • Action/state traces suitable for fine-tuning
  • Task and embodiment metadata for transfer analysis
Dataset signals
ObservationLanguage intentActionsRobot stateEmbodiment
Model hubModel hub
curated source· verification pending· code declared· weights declared
Evaluate in Bench
PolicyOpen2024model hub

RDT-1B

Robotics Diffusion Transformer research

Detail

Diffusion foundation model for bimanual manipulation that uses large-scale robot data to generate action trajectories.

Recommended for

Bimanual manipulation success

Current blocker

Useful for showing why action/state tags need to distinguish single-arm, bimanual, and dexterous data.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Bimanual demonstrations with synchronized visual observations and action/state traces
  • Language or task conditioning for manipulation goals
  • Contact-rich failure and recovery cases for robust long-horizon behavior
Dataset signals
ObservationLanguage intentActionsRobot stateTrajectory
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
PolicyOpen2023model hub

ACT

ALOHA / LeRobot ecosystem

Detail

Action Chunking with Transformers predicts short action sequences for efficient imitation learning in manipulation tasks.

Recommended for

Precision manipulation success

Current blocker

A practical baseline for dataset readiness because it needs clean action supervision.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • High-quality demonstrations with image and state observations
  • Temporally aligned action chunks
  • Task-specific splits that expose compounding-error failures
Dataset signals
ObservationActionsRobot stateHand poseTask success
DocsModel hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
PolicyResearch2023model hub

Diffusion Policy

Columbia / Toyota Research Institute / collaborators

Detail

Visuomotor policy approach that represents robot behavior as a conditional denoising diffusion process.

Recommended for

Multi-modal action generation

Current blocker

Important policy baseline for comparing whether a dataset needs a larger VLA/WAM or a stronger task policy.

Code, weights, and checkpoints
65confidence 40
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Expert demonstrations with action sequences
  • Visual and low-dimensional state observations
  • Failure-aware evaluation splits for multi-modal action distributions
Dataset signals
ObservationActionsRobot stateTrajectoryTask success
Model hub
curated source· verification pending· code unknown· weights declared
Evaluate in Bench
PolicyResearch2023

RoboCat

Google DeepMind

Detail

Self-improving generalist robotic agent that collects new demonstrations to improve its own manipulation capabilities.

Recommended for

Self-improvement loop quality

Current blocker

A useful reference for OpenBot's collect-evaluate-improve loop, even without public weights.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Robot demonstrations connected to self-generated data collection
  • Task success and failure feedback to decide what to collect next
  • Cross-task and cross-robot traces that preserve adaptation history
Dataset signals
ObservationActionsRobot stateTask successFeedback/failure
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAClosed ref2023

RT-2

Google DeepMind

Detail

Vision-language-action model that transfers web-scale vision-language knowledge into robotic control.

Recommended for

Semantic generalization

Current blocker

Useful reference point for VLA direction even if the model itself is not open.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Web-scale semantic knowledge plus robot action data
  • Language-conditioned tasks with visual observations
  • Held-out semantic generalization and physical execution tests
Dataset signals
ObservationLanguage intentActionsObject labelsTask success
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
VLAResearch2023

RT-X

Open X-Embodiment collaboration

Detail

Cross-embodiment RT model family trained from the Open X-Embodiment mixture to study transfer across robots and tasks.

Recommended for

Cross-embodiment transfer

Current blocker

Important for OpenBot because dataset lineage and embodiment metadata become training variables.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Multi-robot trajectories with normalized observations and actions
  • Embodiment metadata that preserves robot, camera, and action-space differences
  • Cross-dataset evaluation to measure whether scaling mixtures improves transfer
Dataset signals
ObservationLanguage intentActionsRobot stateEmbodiment
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
World ModelResearch2023

UniSim

UC Berkeley / research

Detail

Interactive real-world simulator research that models how visual scenes change under actions and interaction.

Recommended for

Interactive scene prediction

Current blocker

Useful reference for OpenBot Synth because generated replay should stay anchored to real interaction traces.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Video traces with action or interaction conditioning
  • Temporal consistency signals and scene geometry
  • Evaluation setups that connect predicted rollouts to downstream control or planning
Dataset signals
ObservationActionsFuture stateCamera poseSim-real
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
PolicyResearch2022

RT-1

Google Robotics

Detail

Robotics Transformer policy trained on large-scale real-world robot demonstrations for language-conditioned manipulation.

Recommended for

Real-world task success

Current blocker

A useful baseline for understanding why policy learning needs action/state alignment, not only video.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Large-scale robot episodes with images, language commands, and actions
  • Task diversity across objects, scenes, and long-horizon instructions
  • Robust train/test splits that expose distribution shift and recovery limits
Dataset signals
ObservationLanguage intentActionsRobot stateTask success
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
PolicyResearch2022

SayCan

Google Research / Everyday Robots

Detail

Language-grounded robotics approach that combines language-model planning with affordance scores from robot skills.

Recommended for

Instruction decomposition

Current blocker

Important for tagging datasets with task phase and feedback, not just final action traces.

Code, weights, and checkpoints
50confidence 10
Loading and training reproducibility
50confidence 15
Training data requirements
76confidence 50
Loop data needed
  • Task instructions decomposed into executable skill phases
  • Affordance or feasibility signals for each skill in the environment
  • Failure labels showing where language plans diverge from physical capability
Dataset signals
ObservationLanguage intentTask phaseFeedback/failureActions
curated source· verification pending· code unknown· weights unknown
Evaluate in Bench
OpenBot connection

The model map turns papers into product requirements.

Catalog scores loop signals. Data preserves them. Bench tests failures. Planned Synth work will later replay measured gaps.

Next useful capability

Link each dataset to policy learning, world models, WAM, and failure-mining value.

Open dataset catalog