TartanAir V2
TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
- Scale
- Not declared
- Formats
- custom
- License
- CC-BY-4.0
- Published
- 2024-03-01
Decision summary
scripted
Metadata requires review against the official source before publication.
Read the dataset manifest and feature schema.
Catalog assessment
Selection evidence
76/100
Provisional · confidence 38 · 3/6 evaluated
Access and governance
Access and license are declared by the source.
Schema and signal coverage
No machine-readable schema has been verified yet.
Policy training readiness
Observation and action/state signals are declared; alignment quality still depends on sample verification.
View 3 more dimensions
World-model readiness
Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.
Failure and recovery readiness
No verified failure/recovery annotation evidence is available yet.
Download and processing readiness
Dataset scale and sampling path are unknown.
Review unresolved evidence and next checks
Signal gaps
- Feedback / correction / failure. Not enough evidence is available to classify this signal.
Next checks
- Read the dataset manifest and feature schema.
- Run a bounded sample audit before assigning Strong readiness.
Metadata coverage
Loop signals
6/7 present or partial
No decision-grade evidence captured yet.
unknownView 6 more signal categories
depth · video · TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
presentTartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
partialTartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
partialTartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
partialdepth · simulation · TartanAir V2 is a large-scale photorealistic synthetic dataset built on Unreal Engine 4 with the AirSim plugin, developed by CMU's AirLab (castacks) to push the limits of visual SLAM, navigation, and robotics perception. It contains 65 highly distinct simulated environments covering urban, rural, domestic, infrastructure, thematic, and nature scenarios, both indoor and outdoor, with some environments split into sub-environments for weather/time-of-day variation. Data is collected from pre-recorded, challenging and realistic trajectories of a generic free-flying 6-DoF camera that mimic real-world robot motion (trajectories are sampled in free space, connected via RRT*, refined for loop closures, and smoothed for physically plausible motion). In each environment 12 perfectly synchronized cameras (two stereo sets of six cameras pointing in six directions to cover a 360-degree view) capture raw pinhole RGB at 640x640, 90-degree FoV, 10 Hz, with a 0.25 m stereo baseline. Raw data is processed into a rich set of modalities: stereo RGB images (plus 1000 Hz MP4 video of the left-front camera), float32 depth maps (compressed losslessly to 4-channel 8-bit PNG), category-level semantic segmentation spanning 1447 semantic classes (manually labeled, with per-environment seg_label_map.json and statistics), optical flow (CUDA-accelerated, generatable across any camera-model pair, stored as npz with covisibility/FoV masks), LiDAR point clouds sampled from depth following Velodyne VLP-16 (and VLP-32C) patterns, IMU with configurable noise modeling, event-camera data generated at 1000 Hz via the ESIM simulator with contrast thresholds 0.2-1.0, occupancy maps, and ground-truth camera poses. The accompanying tartanairpy toolkit supports sampling customizable camera models including pinhole, fisheye (doublesphere / Linear Spherical model), and equirectangular/panoramic views, plus tools for adding noise and motion blur. The dataset is downloaded and managed through the 'tartanair' Python package (pip install tartanair), with data hosted on the AirLab server and Hugging Face (theairlabcmu/tartanair, ~1.11 TB). It succeeds TartanAir V1 (IROS 2020, the official CVPR 2020 Visual SLAM Challenge dataset), adding more scenes, more modalities, and the new fisheye/panoramic camera support; it also underpins the related TartanGround ground-robot dataset.
presentOpen · CC-BY-4.0 · custom
presentEvidence details and provenance
Official signal claims
Schema and annotations
No machine-readable schema facts are captured.
Sample verification
Pending. Metadata does not prove sample coverage, alignment, or file integrity.
curated official source
Integration notes
- Metadata requires review against the official source before publication.
Catalog links
