OpenBot
Back to Explore
Manipulation datasetOpen

TartanAir V2

TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

Scale
Not declared
Formats
custom
License
CC-BY-4.0
Published
2024-03-01

Decision summary

Best for

scripted

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Catalog assessment

Selection evidence

76/100

Provisional · confidence 38 · 3/6 evaluated

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 70 · confidence 55
View 3 more dimensions

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Dataset scale and sampling path are unknown.

unknownnot scored · confidence 10
Review unresolved evidence and next checks

Signal gaps

  • Feedback / correction / failure. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Record specifics

Dataset facts

Source
CMU AirLab
Evidence
secondary claim
Formats
custom
Read paper

Metadata coverage

Loop signals

6/7 present or partial

Feedback / correction / failure

No decision-grade evidence captured yet.

unknown
View 6 more signal categories
Observation / ego video

depth · video · TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

present
Action / hand pose / robot state

TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

partial
Gaze / attention

TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

partial
Language intent / task phase

TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

partial
Sim-real pairing

depth · simulation · TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.

present
License / format / access

Open · CC-BY-4.0 · custom

present
Evidence details and provenance

Official signal claims

DepthEventPoint_cloudProprioceptionSegmentationVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.

Catalog links

Related records