Home › Datasets › task › Autonomous Vehicles

Autonomous Vehicles datasets

archive 2025-07-28

29 datasets carry the task tag "Autonomous Vehicles" (the task itself: Autonomous Vehicles), ordered by the archive's paper count. Page 1 of 1: 29 shown of 29. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Autonomous Vehicles datasets 1–29 of 29

CARLA (Car Learning to Act)
CARLA (CAR Learning to Act) is an open simulator for urban driving, developed as an open-source layer over Unreal Engine 4.
1,345 papers · 4 benchmarks
AirSim is a simulator for drones, cars and more, built on Unreal Engine.
285 papers · 0 benchmarks
The INTERACTION dataset contains naturalistic motions of various traffic participants in a variety of highly interactive driving scenarios from different countries.
81 papers · 1 benchmark
The Talk2Car dataset finds itself at the intersection of various research domains, promoting the development of cross-disciplinary solutions for improving the state-of-the-art in grounding natural language into visual space.
45 papers · 0 benchmarks
The Argoverse 2 Motion Forecasting Dataset is a curated collection of 250,000 scenarios for training and validation.
39 papers · 0 benchmarks
ROAD (ROAD: The ROad event Awareness Dataset for Autonomous Driving)
ROAD is designed to test an autonomous vehicle's ability to detect road events, defined as triplets composed by an active agent, the action(s) it performs and the corresponding scene locations.
27 papers · 0 benchmarks
RadarScenes is a real-world radar point cloud dataset for automotive applications.
27 papers · 0 benchmarks
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
ApolloCar3DT is a dataset that contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints.
17 papers · 14 benchmarks
DADA-2000 is a large-scale benchmark with 2000 video sequences (named as DADA-2000) is contributed with laborious annotation for driver attention (fixation, saccade, focusing time), accident objects/intervals, as well as the accident…
16 papers · 0 benchmarks
TITAN consists of 700 labeled video-clips (with odometry) captured from a moving vehicle on highly interactive urban traffic scenes in Tokyo.
15 papers · 0 benchmarks
PreSIL (Precise Synthetic Image and LiDAR)
Consists of over 50,000 frames and includes high-definition images with full resolution depth information, semantic segmentation (images), point-wise segmentation (point clouds), and detailed annotations for all vehicles and people.
13 papers · 0 benchmarks
comma 2k19 is a dataset of over 33 hours of commute in California's 280 highway.
11 papers · 0 benchmarks
SynthCity is a 367.9M point synthetic full colour Mobile Laser Scanning point cloud.
8 papers · 0 benchmarks
A challenging multi-agent seasonal dataset collected by a fleet of Ford autonomous vehicles at different days and times during 2017-18.
7 papers · 0 benchmarks
A self-driving dataset for motion prediction, containing over 1,000 hours of data.
7 papers · 0 benchmarks
The EuroCity Persons dataset provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes.
6 papers · 0 benchmarks
TCG (Traffic Control Gesture)
The TCG dataset is used to evaluate Traffic Control Gesture recognition for autonomous driving.
4 papers · 1 benchmark
CITR & DUT (CITR dataset and DUT dataset)
Consists of two pedestrian trajectory datasets, CITR dataset and DUT dataset, so that the pedestrian motion models can be further calibrated and verified, especially when vehicle influence on pedestrians plays an important role.
3 papers · 0 benchmarks
HSD (Honda Scenes Dataset)
An annotated dataset is released to enable dynamic scene classification that includes 80 hours of diverse high quality driving video data clips collected in the San Francisco Bay area.
3 papers · 0 benchmarks
TuSimple Lane is an extension of the TuSimple dataset with 14,336 lane boundaries annotations.
3 papers · 0 benchmarks
EyeCar is a dataset of driving videos of vehicles involved in rear-end collisions paired with eye fixation data captured from human subjects.
2 papers · 0 benchmarks
CITR Dataset consists of experimentally designed fundamental VCI scenarios (front, back, and lateral VCIs) and provides unique ID for each pedestrian, which is suitable for exploring a specific aspect of VCI.
1 paper · 0 benchmarks
IN2LAAMA is a set of lidar-inertial datasets collected with a Velodyne VLP-16 lidar and a Xsens MTi-3 IMU.
1 paper · 0 benchmarks
MVX (Multimodal V2X)
MVX incorporates realistic physical world simulation with a differentiable accurate ray tracing wireless simulation that includes multi-agent and multimodal datasets for AI-driven digital twin applications in vehicular communication…
1 paper · 1 benchmark
PRECOG (PREdiction of Clinical Outcomes from Genomic Profiles)
The PREdiction of Clinical Outcomes from Genomic profiles (or PRECOG) encompasses 166 cancer expression data sets, including overall survival data for ~18,000 patients diagnosed with 39 distinct malignancies.
1 paper · 0 benchmarks
Panoramic Video Panoptic Segmentation Dataset is a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving.
1 paper · 0 benchmarks
This dataset is a collection of 4,000 images of cars in multiple scenes that are ready to use for optimizing the accuracy of computer vision models.
0 papers · 0 benchmarks
Hyper Drive (Hyperspectral Driving Dataset)
Towards automated analysis of large environments, hyperspectral sensors must be adapted into a format where they can be operated from mobile robots.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.