SpatialVLA
Shanghai AI Laboratory / collaborators
A spatially enhanced 4B VLA pretrained on 1.1 million real-robot episodes with explicit geometry-aware representations.
Spatial grounding
Official training references Open X-Embodiment and RH20T.
A model hub link is a declaration. Files, loadability, evaluation, and deployment are scored separately.
Model decision scorecard
Use-case scores and evidence confidence are separate; Unknown is not treated as failure.
Code, weights, and checkpoints
UsefulAn official model hub is linked; weight files and loadability are not verified.
Loading and training reproducibility
UnknownNo verified loading configuration is available.
Training data requirements
UsefulRequired signal categories are declared; exact tensor and action interfaces still need verification.
Evaluation evidence
UsefulEvaluation focus is declared, but metrics are not independently verified.
Deployment readiness
UnknownHardware, latency, dependencies, and runtime loading are not yet verified.
Artifact facts and provenance
No metadata-verified artifact facts yet. Source links remain declarations only.
Loop signal demand
Signals this model family needs for training, evaluation, or failure mining.
Observation / ego video
observation · depth · Depth and spatial supervision
Language intent / task phase
language intent
Action / robot state
actions · robot state · RLDS action/state alignment · Efficient action decoding
Future state / dynamics
Needs future-state supervision or rollout structure to validate predictive dynamics.
Feedback / correction / failure
SimplerEnv policy evaluation
Sim-real / embodiment metadata
depth · robot state · Large-scale real robot episodes · Depth and spatial supervision
Evaluation focus
- Spatial grounding
- SimplerEnv policy evaluation
- Efficient action decoding
Missing critical loop signals
Core signal demands are represented. Check quality, alignment, and access constraints.
Related catalog datasets
Open X-Embodiment
Cross-embodiment robot learning in a unified format
Exact action dimensions, control frequency, normalization, and camera mapping require interface verification.
RH20T
Contact-rich multimodal manipulation paired with human demonstrations
Dataset license restricts commercial use.
BridgeData V2
Language- and goal-conditioned manipulation
Exact action dimensions, control frequency, normalization, and camera mapping require interface verification.
OpenBot notes
- Official training references Open X-Embodiment and RH20T.
