OpenBot
Back to Explore
Manipulation datasetOpen

DexYCB

DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.

Scale
0.83 hours
Formats
custom
License
CC-BY-NC-4.0
Published
2021-04-09

Decision summary

Best for

human_demo

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

61/100

Provisional · confidence 41 · 3/6 evaluated

Access and governance

The declared license restricts commercial use.

limitedfit 35 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation/action alignment has not been established.

unknownnot scored · confidence 15
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Feedback / correction / failure. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
NVIDIA (with University of Washington)
Evidence
secondary claim
Formats
custom
episodes
1000
hours
0.83
bytes
127775604736
Read paper

Metadata coverage

Loop signals

5/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
View 5 more signal categories
Observation / ego video

depth · video · DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.

present
Action / hand pose / robot state

DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.

present
Language intent / task phase

DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.

partial
Sim-real pairing

depth · DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.

present
License / format / access

Open · CC-BY-NC-4.0 · custom

partial
Evidence details and provenance

Official signal claims

DepthSegmentationVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records