Home › Datasets › task › Object Detection

Object Detection datasets

archive 2025-07-28

332 datasets carry the task tag "Object Detection" (the task itself: Object Detection), ordered by the archive's paper count. Page 3 of 7: 48 shown of 332. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Object Detection datasets 97–144 of 332

A Multi-Task 4D Radar-Camera Fusion Dataset for Autonomous Driving on Water Surfaces description of the dataset WaterScenes, the first multi-task 4D radar-camera fusion dataset on water surfaces, which offers data from multiple sensors,…
13 papers · 2 benchmarks
GMOT-40 (Generic Multiple Object Tracking (GMOT))
GMOT-40 is the first public dense dataset for Generic Multiple Object Tracking (GMOT).
12 papers · 2 benchmarks
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
People-Art is an object detection dataset which consists of people in 43 different styles.
12 papers · 2 benchmarks
SODA10M is a large-scale object detection benchmark for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw data.
12 papers · 0 benchmarks
TJU-DHD is a high-resolution dataset for object detection and pedestrian detection.
12 papers · 2 benchmarks
The TrashCan dataset is an instance-segmentation dataset of underwater trash.
12 papers · 0 benchmarks
WiderPerson contains a total of 13,382 images with 399,786 annotations, i.e., 29.87 annotations per image, which means this dataset contains dense pedestrians with various kinds of occlusions.
11 papers · 1 benchmark
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
DeepScores contains high quality images of musical scores, partitioned into 300,000 sheets of written music that contain symbols of different shapes and sizes.
10 papers · 0 benchmarks
Description Detection Dataset (D³, /dikju:b/) is an attempt at creating a next-generation object detection dataset.
10 papers · 1 benchmark
GRAZPEDWRI-DX is a public dataset of 20,327 pediatric wrist trauma X-ray images released by the University of Medicine of Graz.
10 papers · 4 benchmarks
SpaceNet 2 (SpaceNet 2: Building Detection v2)
SpaceNet 2: Building Detection v2 - is a dataset for building footprint detection in geographically diverse settings from very high resolution satellite images.
10 papers · 1 benchmark
This is a gun detection dataset with 51K annotated gun images for gun detection and other 51K cropped gun chip images for gun classification collected from a few different sources.
9 papers · 6 benchmarks
PhenoBench (PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain)
The PhenoBench dataset contains multiple image segmentation challenges from the agricultural domain.
9 papers · 0 benchmarks
ReDWeb-S is a large-scale challenging dataset for Salient Object Detection.
9 papers · 0 benchmarks
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
APRICOT is a collection of over 1,000 annotated photographs of printed adversarial patches in public locations.
8 papers · 0 benchmarks
Cops-Ref is a dataset for visual reasoning in context of referring expression comprehension with two main features.
8 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
DUO (Detecting Underwater Objects)
DUO is a dataset for Underwater object detection for robot picking.
8 papers · 1 benchmark
Duke Breast Cancer MRI (Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations)
Breast MRI scans of 922 cancer patients from Duke University, with tumor bounding box annotations, clinical, imaging, and many other features, and more.
8 papers · 0 benchmarks
HS-SOD (HyperSpectral Salient Object Detection Dataset)
HS-SOD is a hyperspectral salient object detection dataset with a collection of 60 hyperspectral images with their respective ground-truth binary images and representative rendered colour images (sRGB).
8 papers · 0 benchmarks
The RIT-18 dataset was built for the semantic segmentation of remote sensing imagery.
8 papers · 0 benchmarks
SKU110K-R is a dataset relabeled with oriented bounding boxes based on SKU110K.
8 papers · 0 benchmarks
SOD (small obstacle detection)
Aiming Detect small obstacles, like lost and found.
8 papers · 2 benchmarks
AGAR (Annotated Germs for Automated Recognition)
The Annotated Germs for Automated Recognition (AGAR) dataset is an image database of microbial colonies cultured on an agar plate.
7 papers · 0 benchmarks
Bamboo Dataset is a mega-scale and information-dense dataset for both classification and detection pre-training.
7 papers · 0 benchmarks
BigDetection is a new large-scale benchmark to build more general and powerful object detection systems.
7 papers · 1 benchmark
Freiburg Groceries is a groceries classification dataset consisting of 5000 images of size 256x256, divided into 25 categories.
7 papers · 0 benchmarks
MobilityAids is a dataset for perception of people and their mobility aids.
7 papers · 0 benchmarks
PFN-PIC (PFN Picking Instructions for Commodities Dataset)
This dataset is a collection of spoken language instructions for a robotic system to pick and place common objects.
7 papers · 0 benchmarks
PIDray is a large-scale dataset which covers various cases in real-world scenarios for prohibited item detection, especially for deliberately hidden items.
7 papers · 0 benchmarks
The PS-Battles dataset is gathered from a large community of image manipulation enthusiasts and provides a basis for media derivation and manipulation detection in the visual domain.
7 papers · 0 benchmarks
YT-BB (YouTube-BoundingBoxes)
YouTube-BoundingBoxes (YT-BB) is a large-scale data set of video URLs with densely-sampled object bounding box annotations.
7 papers · 1 benchmark
BAAI-VANJEE is a dataset for benchmarking and training various computer vision tasks such as 2D/3D object detection and multi-sensor fusion.
6 papers · 0 benchmarks
Comprises about 40,000 images where the most suitable objects for 14 tasks have been annotated.
6 papers · 0 benchmarks
The EuroCity Persons dataset provides a large number of highly diverse, accurate and detailed annotations of pedestrians, cyclists and other riders in urban traffic scenes.
6 papers · 0 benchmarks
FAT (Falling Things)
Falling Things (FAT) is a dataset for advancing the state-of-the-art in object detection and 3D pose estimation in the context of robotics.
6 papers · 0 benchmarks
IIIT-AR-13K is created by manually annotating the bounding boxes of graphical or page objects in publicly available annual reports.
6 papers · 0 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
Lytro Illum is a new light field dataset using a Lytro Illum camera.
6 papers · 0 benchmarks
A large-scale video dataset for MOR in aerial videos.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
Prophesee GEN4 Dataset (Prophesee 1 Megapixel Automotive Detection Dataset)
The dataset is split between train, test and val folders.
6 papers · 0 benchmarks
VEDAI (Vehicle Detection in Aerial Imagery)
VEDAI is a dataset for Vehicle Detection in Aerial Imagery, provided as a tool to benchmark automatic target recognition algorithms in unconstrained environments.
6 papers · 1 benchmark
CBC (Complete Blood Count)
The complete blood count (CBC) dataset contains 360 blood smear images along with their annotation files splitting into Training, Testing, and Validation sets.
5 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.