GM-100 Dataset
Real-robot bimanual evaluation benchmark
A 100-task multi-platform benchmark used by the LingBot-VLA team to compare post-trained VLA policies on real robots.
Real-robot VLA evaluation
Author-created benchmark; reported model results are not independent OpenBot reproduction.
Inspect schema and run a bounded sample audit before committing to the full release.
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Verified facts and provenance
Claims, metadata verification, and sample verification are shown separately.
Unknown — no machine-readable schema facts have been captured.
Unknown — metadata conclusions do not prove sample coverage, alignment, or file integrity.
Declared loop signal coverage
Signals inferred from official metadata; Data pipeline verification is still pending.
Observation / ego video
video
Action / hand pose / robot state
actions · robot state
Gaze / attention
No decision-grade evidence captured yet.
Language intent / task phase
task phase · A 100-task multi-platform benchmark used by the LingBot-VLA team to compare post-trained VLA policies on real robots.
Feedback / correction / failure
Evaluation metadata · Real-robot VLA evaluation · Real-robot bimanual evaluation benchmark
Sim-real pairing
No decision-grade evidence captured yet.
License / format / access
Open · Hugging Face dataset · Evaluation metadata
Model and task fit · OpenBot inference
Has observation, action/state proxy, and task or language context.
fit 85 · confidence 55
Has rich observation and semantic context, but limited geometry/sim-real alignment.
fit 68 · confidence 55
Contains observation, intent, action/state, and feedback-like supervision.
fit 88 · confidence 55
Has failure/evaluation-style labels with action or manipulation context.
fit 82 · confidence 55
Good tasks
Blockers and unresolved evidence
- Gaze / attentionunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
- Sim-real pairingunknownNot enough evidence to classify this signal. Verify metadata or a bounded sample.
Raw dataset signals
OpenBot fit
- Real-robot VLA evaluation
- Cross-platform comparison
- Post-training protocol
Integration notes
- Author-created benchmark; reported model results are not independent OpenBot reproduction.
