Home › Datasets › task › Object Detection

Object Detection datasets

archive 2025-07-28

332 datasets carry the task tag "Object Detection" (the task itself: Object Detection), ordered by the archive's paper count. Page 2 of 7: 48 shown of 332. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Object Detection datasets 49–96 of 332

A2D (Actor-Action Dataset)
A2D (Actor-Action Dataset) is a dataset for simultaneously inferring actors and actions in videos.
42 papers · 1 benchmark
H3D (Honda Research Institute 3D)
The H3D is a large scale full-surround 3D multi-object detection and tracking dataset.
39 papers · 0 benchmarks
The SIXray dataset is constructed by the Pattern Recognition and Intelligent System Development Laboratory, University of Chinese Academy of Sciences.
39 papers · 1 benchmark
The A3D dataset is a step forward to make autonomous driving safer for pedestrians and the public in the real world.
37 papers · 0 benchmarks
Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes.
37 papers · 1 benchmark
Watercolor2k is a dataset used for cross-domain object detection which contains 2k watercolor images with image and instance-level annotations.
37 papers · 3 benchmarks
JTA (Joint Track Auto)
JTA is a dataset for people tracking in urban scenarios by exploiting a photorealistic videogame.
36 papers · 1 benchmark
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
34 papers · 2 benchmarks
The EgoHands dataset contains 48 Google Glass videos of complex, first-person interactions between two people.
34 papers · 0 benchmarks
InteriorNet is a RGB-D for large scale interior scene understanding and mapping.
30 papers · 0 benchmarks
AVD (Active Vision Dataset)
AVD focuses on simulating robotic vision tasks in everyday indoor environments using real imagery.
29 papers · 1 benchmark
Comic2k is a dataset used for cross-domain object detection which contains 2k comic images with image and instance-level annotations.
29 papers · 4 benchmarks
Includes 5000 spatially aligned RGBT image pairs with ground truth annotations.
29 papers · 0 benchmarks
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
Object detection benchmark for logo detection.
26 papers · 3 benchmarks
ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer)
The ELEVATER benchmark is a collection of resources for training, evaluating, and analyzing language-image models on image classification and object detection.
25 papers · 2 benchmarks
TinyPerson is a benchmark for tiny object detection in a long distance and with massive backgrounds.
25 papers · 0 benchmarks
CCPD (Chinese City Parking Dataset)
The Chinese City Parking Dataset (CCPD) is a dataset for license plate detection and recognition.
24 papers · 0 benchmarks
The INRIA Person dataset is a dataset of images of persons used for pedestrian detection.
24 papers · 0 benchmarks
RADIATE (RAdar Dataset In Adverse weaThEr)
RADIATE (RAdar Dataset In Adverse weaThEr) is new automotive dataset created by Heriot-Watt University which includes Radar, Lidar, Stereo Camera and GPS/IMU.
24 papers · 2 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
CADC (Canadian Adverse Driving Conditions)
Collected with the Autonomoose autonomous vehicle platform, based on a modified Lincoln MKZ.
23 papers · 0 benchmarks
DSEC (A Stereo Event Camera Dataset for Driving Scenarios)
DSEC is a stereo camera dataset in driving scenarios that contains data from two monochrome event cameras and two global shutter color cameras in favorable and challenging illumination conditions.
23 papers · 2 benchmarks
OpenImages V6 is a large-scale dataset , consists of 9 million training images, 41,620 validation samples, and 125,456 test samples.
23 papers · 2 benchmarks
COWC (Cars Overhead With Context)
The Cars Overhead With Context (COWC) data set is a large set of annotated cars from overhead.
22 papers · 0 benchmarks
IP102 contains more than 75,000 images belonging to 102 categories, which exhibit a natural long-tailed distribution.
22 papers · 0 benchmarks
MINOS is a simulator designed to support the development of multisensory models for goal-directed navigation in complex indoor environments.
22 papers · 0 benchmarks
The Sku110k dataset provides 11,762 images with more than 1.7 million annotated bounding boxes captured in densely packed scenarios, including 8,233 images for training, 588 images for validation, and 2,941 images for testing.
22 papers · 1 benchmark
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
A composite dataset that unifies semantic segmentation datasets from different domains.
21 papers · 0 benchmarks
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
20 papers · 2 benchmarks
EORSSD (Extended Optical Remote Sensing Saliency Detection)
The Extended Optical Remote Sensing Saliency Detection (EORSSD) dataset is an extension of the ORSSD dataset.
20 papers · 0 benchmarks
PlantDoc is a dataset for visual plant disease detection.
20 papers · 1 benchmark
Argoverse-HD is a dataset built for streaming object detection, which encompasses real-time object detection, video object detection, tracking, and short-term forecasting.
19 papers · 4 benchmarks
ModaNet is a street fashion images dataset consisting of annotations related to RGB images.
19 papers · 1 benchmark
AI-TOD (Tiny Object Detection in Aerial Images)
AI-TOD comes with 700,621 object instances for eight categories across 28,036 aerial images.
18 papers · 2 benchmarks
CoIR (Code Information Retrieval Benchmark)
CoIR (Code Information Retrieval) benchmark, is designed to evaluate code retrieval capabilities.
18 papers · 1 benchmark
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
RPC (Retail Product Checkout)
RPC is a large-scale retail product checkout dataset and collects 200 retail SKUs.
17 papers · 0 benchmarks
MinneApple is a benchmark dataset for apple detection and segmentation.
16 papers · 0 benchmarks
MALF (Multi-Attribute Labelled Faces)
The MALF dataset is a large dataset with 5,250 images annotated with multiple facial attributes and it is specifically constructed for fine grained evaluation.
15 papers · 0 benchmarks
Washington RGB-D is a widely used testbed in the robotic community, consisting of 41,877 RGB-D images organized into 300 instances divided in 51 classes of common indoor objects (e.g.
15 papers · 0 benchmarks
A novel dataset for traffic accidents analysis.
14 papers · 0 benchmarks
GEN1 Detection (Prophesee GEN1 Automotive Detection Dataset)
Prophesee’s GEN1 Automotive Detection Dataset is the largest Event-Based Dataset to date.
13 papers · 1 benchmark
PreSIL (Precise Synthetic Image and LiDAR)
Consists of over 50,000 frames and includes high-definition images with full resolution depth information, semantic segmentation (images), point-wise segmentation (point clouds), and detailed annotations for all vehicles and people.
13 papers · 0 benchmarks
TTPLA (Transmission Towers and Power Lines (TTPLA))
TTPLA is a public dataset which is a collection of aerial images on Transmission Towers (TTs) and Power Lines (PLs).
13 papers · 0 benchmarks
UFDD (Unconstrained Face Detection Dataset)
Unconstrained Face Detection Dataset (UFDD) aims to fuel further research in unconstrained face detection.
13 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.