Home › Datasets › task › Object Detection

Object Detection datasets

archive 2025-07-28

332 datasets carry the task tag "Object Detection" (the task itself: Object Detection), ordered by the archive's paper count. Page 1 of 7: 48 shown of 332. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Object Detection datasets 1–48 of 332

The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
Visual Genome contains Visual Question Answering data in a multi-choice setting.
1,256 papers · 15 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
The Waymo Open Dataset is comprised of high resolution sensor data collected by autonomous vehicles operated by the Waymo Driver in a wide variety of conditions.
481 papers · 16 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
Manga109 has been compiled by the Aizawa Yamasaki Matsui Laboratory, Department of Information and Communication Engineering, the Graduate School of Information Science and Technology, the University of Tokyo.
300 papers · 12 benchmarks
Description: 10,000 People - Human Pose Recognition Data.
265 papers · 1 benchmark
Foggy Cityscapes is a synthetic foggy dataset which simulates fog on real scenes.
249 papers · 7 benchmarks
The Pascal3D+ multi-view dataset consists of images in the wild, i.e., images of object categories exhibiting high variability, captured under uncontrolled settings, in cluttered scenes and under many different poses.
237 papers · 1 benchmark
Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
PASCAL VOC (PASCAL Visual Object Classes Challenge)
The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa,…
198 papers · 18 benchmarks
LabelMe database is a large collection of images with ground truth labels for object detection and recognition.
178 papers · 1 benchmark
The nocaps benchmark consists of 166,100 human-generated captions describing 15,100 images from the OpenImages validation and test sets.
175 papers · 13 benchmarks
CrowdHuman is a large and rich-annotated human detection dataset, which contains 15,000, 4,370 and 5,000 images collected from the Internet for training, validation and testing respectively.
161 papers · 2 benchmarks
Objects365 is a large-scale object detection dataset, Objects365, which has 365 object categories over 600K training images.
161 papers · 2 benchmarks
fMoW (Functional Map of the World)
Functional Map of the World (fMoW) is a dataset that aims to inspire the development of machine learning models capable of predicting the functional purpose of buildings and land use from temporal sequences of satellite images and a rich…
144 papers · 1 benchmark
The CityPersons dataset is a subset of Cityscapes which only consists of person annotations.
129 papers · 2 benchmarks
PubLayNet is a dataset for document layout analysis by automatically matching the XML representations and the content of over 1 million PDF articles that are publicly available on PubMed Central.
123 papers · 1 benchmark
Kvasir (The Kvasir Dataset)
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
LLVIP (A Visible-infrared Paired Dataset for Low-light Vision)
Visible-infrared Paired Dataset for Low-light Vision 30976 images (15488 pairs) 24 dark scenes, 2 daytime scenes Support for image-to-image translation (visible to infrared, or infrared to visible), visible and infrared image fusion,…
116 papers · 6 benchmarks
The ABC Dataset is a collection of one million Computer-Aided Design (CAD) models for research of geometric deep learning methods and applications.
104 papers · 0 benchmarks
UAVDT (Unmanned Aerial Vehicle Benchmark Object Detection and Tracking)
UAVDT is a large scale challenging UAV Detection and Tracking benchmark (i.e., about 80, 000 representative frames from 10 hours raw videos) for 3 important fundamental tasks, i.e., object DETection (DET), Single Object Tracking (SOT) and…
96 papers · 2 benchmarks
xView is one of the largest publicly available datasets of overhead imagery.
93 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
90 papers · 13 benchmarks
UIEB (Underwater Image Enhancement Benchmark Dataset)
Includes 950 real-world underwater images, 890 of which have the corresponding reference images.
87 papers · 1 benchmark
A large dataset of human hand images (dorsal and palmar sides) with detailed ground-truth information for gender recognition and biometric identification.
82 papers · 0 benchmarks
iSAID contains 655,451 object instances for 15 categories across 2,806 high-resolution images.
81 papers · 4 benchmarks
ApolloScape is a large dataset consisting of over 140,000 video frames (73 street scene videos) from various locations in China under varying weather conditions.
74 papers · 4 benchmarks
VisDrone is a large-scale benchmark with carefully annotated ground-truth for various important computer vision tasks, to make vision meet drones.
73 papers · 1 benchmark
SPair-71k contains 70,958 image pairs with diverse variations in viewpoint and scale.
71 papers · 2 benchmarks
FSOD (Few-Shot Object Detection Dataset)
Few-Shot Object Detection Dataset (FSOD) is a high-diverse dataset specifically designed for few-shot object detection and intrinsically designed to evaluate thegenerality of a model on novel categories.
68 papers · 0 benchmarks
PASCAL-Part is a set of additional annotations for PASCAL VOC 2010.
65 papers · 4 benchmarks
Contains 51,583 descriptions of 11,046 objects from 800 ScanNet scenes.
63 papers · 1 benchmark
ExDark (Exclusively Dark Image Dataset)
The Exclusively Dark (ExDARK) dataset is a collection of 7,363 low-light images from very low-light environments to twilight (i.e 10 different conditions) with 12 object classes (similar to PASCAL VOC) annotated on both image class level…
58 papers · 2 benchmarks
Fisheye cameras are commonly employed for obtaining a large field of view in surveillance, augmented reality and in particular automotive applications.
55 papers · 1 benchmark
DAQUAR (DAtaset for QUestion Answering on Real-world images) is a dataset of human question answer pairs about images.
54 papers · 0 benchmarks
Consists of 100 challenging video sequences captured from real-world traffic scenes (over 140,000 frames with rich annotations, including occlusion, weather, vehicle category, truncation, and vehicle bounding boxes) for object detection,…
53 papers · 2 benchmarks
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
In Clipart1k, the target domain classes to be detected are the same as those in the source domain.
48 papers · 3 benchmarks
Synscapes is a synthetic dataset for street scene parsing created using photorealistic rendering techniques, and show state-of-the-art results for training and validation as well as new types of analysis.
46 papers · 1 benchmark
COCO-O(ut-of-distribution) contains 6 domains (sketch, cartoon, painting, weather, handmake, tattoo) of COCO objects which are hard to be detected by most existing detectors.
44 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.