BridgeData V2 Dataset
Language- and goal-conditioned manipulation
A diverse WidowX manipulation dataset with natural-language instructions and broad variation in environments, objects, and camera poses.
Open-vocabulary policies
A core public baseline for VLA and policy post-training.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video · depth · A diverse WidowX manipulation dataset with natural-language instructions and broad variation in environments, objects, and camera poses.
Action / hand pose / robot state
actions · robot state · Language- and goal-conditioned manipulation · A diverse WidowX manipulation dataset with natural-language instructions and broad variation in environments, objects, and camera poses.
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
language · Language- and goal-conditioned manipulation · A diverse WidowX manipulation dataset with natural-language instructions and broad variation in environments, objects, and camera poses.
Feedback / correction / failure
No decision-grade evidence captured yet.
Sim-real pairing
depth · A diverse WidowX manipulation dataset with natural-language instructions and broad variation in environments, objects, and camera poses.
License / format / access
Open · RLDS
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has visual observations plus geometry, calibration, depth, or reconstruction cues.
fit 82 · confidence 55
Contains observation, intent, and action/state, but feedback/correction signal is weak.
fit 72 · confidence 55
Failure, correction, intervention, and recovery annotations have not been verified.
fit 50 · confidence 15
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Feedback / correction / failureunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Open-vocabulary policies
- Goal-conditioned learning
- Cross-environment generalization
Related models and papers
Model references linked to similar loop signals.
LingBot-VA
An autoregressive video-action world-model policy that interleaves future video-latent prediction with robot action generation.
pi0.5
A VLA model designed for open-world generalization by combining heterogeneous robot data, semantic subtask prediction, and high-level knowledge transfer.
SpatialVLA
A spatially enhanced 4B VLA pretrained on 1.1 million real-robot episodes with explicit geometry-aware representations.
OpenVLA-OFT
An optimized OpenVLA fine-tuning recipe with continuous action chunks, multi-image input, and faster high-frequency control.
Integration notes
- A core public baseline for VLA and policy post-training.
