DexYCB
DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.
- Scale
- 0.83 hours
- Formats
- custom
- License
- CC-BY-NC-4.0
- Published
- 2021-04-09
Decision summary
human_demo
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
61/100
Provisional · confidence 41 · 3/6 evaluated
Access and governance
The declared license restricts commercial use.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation/action alignment has not been established.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Scale is declared; transfer and processing estimates are not measured.
Review unresolved evidence and next checks
Signal gaps
- Gaze / attention. Not enough evidence is available to classify this signal.
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Record specifics
Dataset facts
- Source
- NVIDIA (with University of Washington)
- Evidence
- secondary claim
- Formats
- custom
- episodes
- 1000
- hours
- 0.83
- bytes
- 127775604736
Metadata coverage
Loop signals
5/7 present or partial
No decision-grade evidence captured yet.
unknownNo decision-grade evidence captured yet.
unknownView 5 more signal categories
depth · video · DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.
presentDexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.
presentDexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.
partialdepth · DexYCB is a large-scale RGB-D dataset of 582,000 frames capturing 10 human subjects grasping 20 objects from the YCB-Video set, recorded synchronously from 8 calibrated Intel RealSense D415 cameras around a tabletop (640x480 @ 30fps, 1,000 trials total). Each frame is annotated with segmentation masks, 6D object poses, MANO hand pose, and 2D/3D hand joints. It supports tasks including 2D detection, 6D object pose estimation, 3D hand pose estimation, and a robotics-relevant safe human-to-robot handover grasp-generation task.
presentOpen · CC-BY-NC-4.0 · custom
partialEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
