OpenBot
Back to Explore
Manipulation datasetOpen

RoboArena

**RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

Scale
10783 episodes
Formats
custom
License
MIT
Published
2025-08-05

Decision summary

Best for

policy_rollout

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

73/100

Provisional · confidence 48 · 4/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 85 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 68 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
RoboArena (multi-institution consortium)
Evidence
secondary claim
Formats
custom
episodes
10783
bytes
21669947196
Read paper

Metadata coverage

Loop signals

6/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

video · **RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

present
Action / hand pose / robot state

**RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

partial
Language intent / task phase

language · has-success-labels · **RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

present
Feedback / correction / failure

**RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

present
Sim-real pairing

**RoboArena** is a living, distributed real-robot benchmark: Chatbot Arena, but the contestants are robot policies and the arena is a network of physical Franka arms. Crucially, **every episode is an autonomous policy rollout, not a teleoperated demonstration** — a human is present only to reset the scene and score the result. - **10,783 policy episodes** across **3,883 evaluation sessions** in the latest snapshot (**21.7 GB**, 27,148 MP4 videos), up from 4,613 episodes in the first public dump. - Every site runs the **DROID platform**: a **Franka Panda** 7-DoF arm with a Robotiq 2F-85 gripper, a ZED-mini stereo **wrist camera**, and ZED 2 stereo **shoulder cameras** (left and right) on a mobile height-adjustable table. - Evaluators at **7+ academic institutions** — Berkeley, Stanford, UW, Montréal, NVIDIA, Penn, UT Austin, Yonsei — each run **pairs of competing policies** on tasks of their own choosing, then rate which did better. There is no fixed task list and no central lab. - **21 policies** appear in the latest index, including pi0, pi0-fast, pi0.5, several PaliGemma variants, MolmoAct2-DROID and Cosmos3-Nano-Policy, spanning both joint-velocity and joint-position action spaces. - Each session records a natural-language task instruction, a **binary A/B/TIE preference**, a **0–100 continuous partial-success score per policy**, a binary success flag, episode duration, and free-form evaluator feedback. - Released as periodic **DataDump** snapshots on Hugging Face (Aug 2025, Feb 2026, Jul 2026), so the corpus grows over time rather than freezing. - Published at **CoRL 2025**. Because it captures how policies actually fail in the wild, it is useful well beyond leaderboards — for failure analysis, reward modelling, and preference learning.

partial
License / format / access

Open · MIT · custom

partial
Evidence details and provenance

Official signal claims

LanguageProprioceptionVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records