OpenBot
Egocentric datasetOpen
OBRS 32Bronze

TartanAir

TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020.

Source notes

TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020. Data is collected in 30 photo-realistic Unreal Engine environments via the AirSim plugin, spanning urban, rural, nature, domestic, public and sci-fi scenes, with challenging conditions including day-night/lighting changes, weather (rain, snow, fog, wind), seasonal variation, moving/dynamic objects, and aggressive diverse ego-motion. It comprises 1037 long motion sequences (each 500-4000 frames) totaling over one million frames, organized hierarchically by Environment -> Difficulty (Easy/Hard) -> Trajectory (P000, P001, ...) -> modality subfolders. Each frame provides multi-modal sensor data and precise ground truth: stereo RGB (left/right, ~640x480), depth maps, semantic segmentation, optical flow (with occlusion masks), 6-DoF camera poses, simulated multi-line LiDAR point clouds, and simulated IMU. Modalities are stored in open formats: RGB as PNG, depth/segmentation/optical flow as NumPy .npy arrays, and poses as .txt files. The full dataset is up to ~3TB and is distributed via the Microsoft Azure Open Datasets platform, an AirLab Ceph/S3 endpoint, and Hugging Face, with download scripts (download_training.py using boto3) and tooling provided at github.com/castacks/tartanair_tools. It served as the official dataset of the CVPR 2020 Visual SLAM Challenge (monocular and stereo tracks). The dataset is released under CC-BY-4.0; the accompanying tartanair_tools software is separately BSD-3-Clause licensed.

Scale
27.8 hours
Formats
custom
License
CC-BY-4.0
Published
2020-03-31

Decision summary

Best for

scripted

Main blocker

Metadata requires review against the official source before publication.

Next check

Read the dataset manifest and feature schema.

Release history

Source-backed release timing for this canonical dataset record.

Important releases
  1. Dataset series published

    Release timing is recorded from official dataset metadata.

    Release evidence

Selection readiness

6 evidence dimensions for deciding whether this dataset is ready to inspect, compare, or adopt. This is not a model benchmark.

Catalog evidence · not task performance

73/100

Provisional

4/6 dimensions scored · 48% confidence

Access and governance75conf. 85
Schema and signal coverageNot scoredconf. 15
Policy training readiness70conf. 55
World-model readiness82conf. 50
Failure and recovery readinessNot scoredconf. 15
Download and processing readiness65conf. 65
Read the evidence behind all 6 dimensions

Access and governance

Access and license are declared by the source.

usefulfit 75 · confidence 85

Schema and signal coverage

No machine-readable schema has been verified yet.

unknownnot scored · confidence 15

Policy training readiness

Observation and action/state signals are declared; alignment quality still depends on sample verification.

usefulfit 70 · confidence 55

World-model readiness

Temporal observations plus geometry or semantic context are declared; sample alignment remains to be audited.

usefulfit 82 · confidence 50

Failure and recovery readiness

No verified failure/recovery annotation evidence is available yet.

unknownnot scored · confidence 15

Download and processing readiness

Scale is declared; transfer and processing estimates are not measured.

usefulfit 65 · confidence 65
Review unresolved evidence and next checks

Signal gaps

  • Gaze / attention. Not enough evidence is available to classify this signal.
  • Language intent / task phase. Not enough evidence is available to classify this signal.

Next checks

  • Read the dataset manifest and feature schema.
  • Run a bounded sample audit before assigning Strong readiness.

Dataset facts

Source
CMU AirLab
Evidence
secondary claim
Formats
custom
episodes
1,037
hours
27.8
size
3.00 TB
Read paper

Loop signals

5/7 present or partial

Gaze / attention

No decision-grade evidence captured yet.

unknown
Language intent / task phase

No decision-grade evidence captured yet.

unknown
View 5 more signal categories
Observation / ego video

depth · video · TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020. Data is collected in 30 photo-realistic Unreal Engine environments via the AirSim plugin, spanning urban, rural, nature, domestic, public and sci-fi scenes, with challenging conditions including day-night/lighting changes, weather (rain, snow, fog, wind), seasonal variation, moving/dynamic objects, and aggressive diverse ego-motion. It comprises 1037 long motion sequences (each 500-4000 frames) totaling over one million frames, organized hierarchically by Environment -> Difficulty (Easy/Hard) -> Trajectory (P000, P001, ...) -> modality subfolders. Each frame provides multi-modal sensor data and precise ground truth: stereo RGB (left/right, ~640x480), depth maps, semantic segmentation, optical flow (with occlusion masks), 6-DoF camera poses, simulated multi-line LiDAR point clouds, and simulated IMU. Modalities are stored in open formats: RGB as PNG, depth/segmentation/optical flow as NumPy .npy arrays, and poses as .txt files. The full dataset is up to ~3TB and is distributed via the Microsoft Azure Open Datasets platform, an AirLab Ceph/S3 endpoint, and Hugging Face, with download scripts (download_training.py using boto3) and tooling provided at github.com/castacks/tartanair_tools. It served as the official dataset of the CVPR 2020 Visual SLAM Challenge (monocular and stereo tracks). The dataset is released under CC-BY-4.0; the accompanying tartanair_tools software is separately BSD-3-Clause licensed.

present
Action / hand pose / robot state

TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020. Data is collected in 30 photo-realistic Unreal Engine environments via the AirSim plugin, spanning urban, rural, nature, domestic, public and sci-fi scenes, with challenging conditions including day-night/lighting changes, weather (rain, snow, fog, wind), seasonal variation, moving/dynamic objects, and aggressive diverse ego-motion. It comprises 1037 long motion sequences (each 500-4000 frames) totaling over one million frames, organized hierarchically by Environment -> Difficulty (Easy/Hard) -> Trajectory (P000, P001, ...) -> modality subfolders. Each frame provides multi-modal sensor data and precise ground truth: stereo RGB (left/right, ~640x480), depth maps, semantic segmentation, optical flow (with occlusion masks), 6-DoF camera poses, simulated multi-line LiDAR point clouds, and simulated IMU. Modalities are stored in open formats: RGB as PNG, depth/segmentation/optical flow as NumPy .npy arrays, and poses as .txt files. The full dataset is up to ~3TB and is distributed via the Microsoft Azure Open Datasets platform, an AirLab Ceph/S3 endpoint, and Hugging Face, with download scripts (download_training.py using boto3) and tooling provided at github.com/castacks/tartanair_tools. It served as the official dataset of the CVPR 2020 Visual SLAM Challenge (monocular and stereo tracks). The dataset is released under CC-BY-4.0; the accompanying tartanair_tools software is separately BSD-3-Clause licensed.

present
Feedback / correction / failure

TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020. Data is collected in 30 photo-realistic Unreal Engine environments via the AirSim plugin, spanning urban, rural, nature, domestic, public and sci-fi scenes, with challenging conditions including day-night/lighting changes, weather (rain, snow, fog, wind), seasonal variation, moving/dynamic objects, and aggressive diverse ego-motion. It comprises 1037 long motion sequences (each 500-4000 frames) totaling over one million frames, organized hierarchically by Environment -> Difficulty (Easy/Hard) -> Trajectory (P000, P001, ...) -> modality subfolders. Each frame provides multi-modal sensor data and precise ground truth: stereo RGB (left/right, ~640x480), depth maps, semantic segmentation, optical flow (with occlusion masks), 6-DoF camera poses, simulated multi-line LiDAR point clouds, and simulated IMU. Modalities are stored in open formats: RGB as PNG, depth/segmentation/optical flow as NumPy .npy arrays, and poses as .txt files. The full dataset is up to ~3TB and is distributed via the Microsoft Azure Open Datasets platform, an AirLab Ceph/S3 endpoint, and Hugging Face, with download scripts (download_training.py using boto3) and tooling provided at github.com/castacks/tartanair_tools. It served as the official dataset of the CVPR 2020 Visual SLAM Challenge (monocular and stereo tracks). The dataset is released under CC-BY-4.0; the accompanying tartanair_tools software is separately BSD-3-Clause licensed.

partial
Sim-real pairing

depth · simulation · TartanAir is a large-scale synthetic dataset for visual SLAM and robot navigation, released by CMU's AirLab (Robotics Institute) and presented at IROS 2020. Data is collected in 30 photo-realistic Unreal Engine environments via the AirSim plugin, spanning urban, rural, nature, domestic, public and sci-fi scenes, with challenging conditions including day-night/lighting changes, weather (rain, snow, fog, wind), seasonal variation, moving/dynamic objects, and aggressive diverse ego-motion. It comprises 1037 long motion sequences (each 500-4000 frames) totaling over one million frames, organized hierarchically by Environment -> Difficulty (Easy/Hard) -> Trajectory (P000, P001, ...) -> modality subfolders. Each frame provides multi-modal sensor data and precise ground truth: stereo RGB (left/right, ~640x480), depth maps, semantic segmentation, optical flow (with occlusion masks), 6-DoF camera poses, simulated multi-line LiDAR point clouds, and simulated IMU. Modalities are stored in open formats: RGB as PNG, depth/segmentation/optical flow as NumPy .npy arrays, and poses as .txt files. The full dataset is up to ~3TB and is distributed via the Microsoft Azure Open Datasets platform, an AirLab Ceph/S3 endpoint, and Hugging Face, with download scripts (download_training.py using boto3) and tooling provided at github.com/castacks/tartanair_tools. It served as the official dataset of the CVPR 2020 Visual SLAM Challenge (monocular and stereo tracks). The dataset is released under CC-BY-4.0; the accompanying tartanair_tools software is separately BSD-3-Clause licensed.

present
License / format / access

Open · CC-BY-4.0 · custom

present
Evidence details and provenance

Official signal claims

DepthPoint_cloudProprioceptionSegmentationVideo

Schema and annotations

No machine-readable schema facts are captured.

Sample verification

Pending. Metadata does not prove sample coverage, alignment, or file integrity.

curated official source

Integration notes

  • Metadata requires review against the official source before publication.