Humanoid VLA training
Humanoid VLA training
Non-commercial share-alike license; commercial use requires separate legal review or permission.
Compare access, license, modality, signal, and scale before a training run.
Showing 39 of 39 datasets · 26 open · 13 restricted · 22 on Hugging Face · published catalog
Humanoid VLA training
Humanoid VLA training
Non-commercial share-alike license; commercial use requires separate legal review or permission.
Open-vocabulary policies
Open-vocabulary policies
A core public baseline for VLA and policy post-training.
VLA fine-tuning
VLA fine-tuning
Strong real-world diversity; verify release subset completeness before processing.
LeRobot conversion tests
LeRobot conversion tests
One of the more directly robot-task-shaped open HF entries.
Embodied reasoning
Embodied reasoning
Separate EO-Data1.5M training samples from EO-Bench evaluation data.
Mobile manipulation
Mobile manipulation
Strong source for connecting subtask annotation to mobile manipulation evaluation.
Real-robot VLA evaluation
Real-robot VLA evaluation
Author-created benchmark; reported model results are not independent OpenBot reproduction.
Open VLA pretraining
Open VLA pretraining
Licenses and annotation quality vary by repository; treat this as a collection with per-source governance.
Long-horizon policy learning
Long-horizon policy learning
Track as a LIBERO dataset release, not as an unrelated top-level series once versioned Catalog storage is active.
Household skill segmentation
Household skill segmentation
One of the more directly robotics-relevant egocentric datasets in this catalog.
Cross-embodiment VLA pretraining
Cross-embodiment VLA pretraining
Licenses differ across constituent datasets and must be checked individually.
Failure mining
Failure mining
One of the strongest current matches for OpenBot loop-signal analysis.
Bimanual VLA post-training
Bimanual VLA post-training
This entry covers a generator, released trajectories, and benchmark; keep those artifacts distinct in the future release schema.
World model pretraining
World model pretraining
Very large and controlled-access; the practical OpenBot path is metadata indexing plus targeted subset pulls.
Learning from skilled human demonstrations
Learning from skilled human demonstrations
Especially relevant when a task needs both wearable camera context and external validation views.
Pick-and-place benchmark fixtures
Pick-and-place benchmark fixtures
Open access but non-commercial license terms apply.
Schema validation for LeRobot v3
Schema validation for LeRobot v3
Small, but excellent for testing catalog pages, loaders, and schema conversions.
Kitchen manipulation vocabulary
Kitchen manipulation vocabulary
Still a strong reference point for kitchen-centric manipulation data.
Action-label schema tests
Action-label schema tests
Preview-scale dataset, not the full corpus.
Testing LeRobot import paths
Testing LeRobot import paths
Good open sample for OpenBot Data loaders and dataset pages.
Dexterous-hand data demos
Dexterous-hand data demos
Open access on Hugging Face, but non-commercial license terms apply.
Contact-rich policy learning
Contact-rich policy learning
Commercial-use rights differ by subset; privacy-sensitive human recordings are included.
View synthesis for wrist/ego cameras
View synthesis for wrist/ego cameras
Short recordings, but high camera density and strong reconstruction value.
Fast catalog smoke tests
Fast catalog smoke tests
Small but concrete; useful for demos and ingestion tests.
Pretraining egocentric perception
Pretraining egocentric perception
Best treated as a large source corpus rather than a direct robot-action dataset.
Activity-recognition examples
Activity-recognition examples
Good open sample for testing language annotations around egocentric clips.
Affordance pretraining
Affordance pretraining
Useful for dataset curation labels even when source videos are not robot trajectories.
Comparing egocentric corpora before ingest
Comparing egocentric corpora before ingest
Open evaluation layer; the larger source datasets may still be gated.
Industrial manipulation priors
Industrial manipulation priors
Good candidate for OpenBot Data indexing, dedup, and active-manipulation scoring.
Benchmarking egocentric data quality
Benchmarking egocentric data quality
Good open-access companion to the gated Egocentric-10K source dataset.
Wearable-assistant evaluation
Wearable-assistant evaluation
A practical bridge between video-only egocentric data and robot-ready sensor streams.
Video-language model evaluation
Video-language model evaluation
Not a manipulation training set, but helpful for evaluating whether agents understand long first-person context.
Dataset registry examples
Dataset registry examples
Best treated as a registry source rather than raw training data.
Object persistence in robot tasks
Object persistence in robot tasks
Useful for Bench/Data integration when failures involve losing an object through a manipulation step.
Residential manipulation examples
Residential manipulation examples
Small open-access sample, useful for demos rather than model-scale training.
Industrial manipulation examples
Industrial manipulation examples
Useful for examples and smoke tests; scale is intentionally small.
Depth completion
Depth completion
Perception dataset rather than an action-policy training corpus; RobbyVla is the manipulation-related subset.
Small open fixture for first-person video ingestion
Small open fixture for first-person video ingestion
Small sample dataset, not a large training corpus.
Agent evaluation over streaming first-person video
Agent evaluation over streaming first-person video
Less robot-action focused, but useful for evaluating an agent's temporal grounding.
OpenBot Data turns raw sources into versioned, replay-ready training data.