Home › Datasets › task › Object Detection
Object Detection datasets
archive 2025-07-28
332 datasets carry the task tag "Object Detection" (the task itself: Object Detection), ordered by the archive's paper count. Page 6 of 7: 48 shown of 332. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Object Detection datasets 241–288 of 332
HGP (Hands Guns and Phones Dataset)
Hands Guns and Phones (HGP) dataset contains 2199 images (1989 for training an 210 for testing) of people using guns or phones in real-world scenarios (people making phones reviews, shooting drills, or making calls).
1 paper · 0 benchmarks
We establish the first large benchmark called IRBFD to facilitate the research in the area of nonuniformity correction and infrared UAV target detection, which consists of 50,000 manually labeled infrared images with various nonuniformity…
1 paper · 0 benchmarks
IndraEye (IndraEye: Infrared Electro-Optical Drone-based Aerial Object Detection Dataset)
Deep neural networks (DNNs) have demonstrated superior performance when trained on well-illuminated environments, given that the images are captured through an Electro-Optical (EO) camera, which offers rich texture content.
1 paper · 0 benchmarks
LDD (LDD: A Grape Diseases Dataset Detection and Instance Segmentation)
The Instance Segmentation task, an extension of the well-known Object Detection task, is of great help in many areas, such as precision agriculture: being able to automatically identify plant organs and the possible diseases associated…
1 paper · 2 benchmarks
LVVO (Lecture Video Visual Objects)
The Lecture Video Visual Objects (LVVO) dataset is a benchmark designed for object detection in lecture video frames.
1 paper · 0 benchmarks
LiDAR-CS is a dataset for 3D object detection in real traffic.
1 paper · 0 benchmarks
Loucount is a retail object detection and and counting dataset with rich annotations in retail stores, which consists of 50, 394 images with more than 1.9 million object instances in 140 categories
1 paper · 0 benchmarks
A small-scale training set, which only contains 4K images.
1 paper · 0 benchmarks
METU-ALET is an image dataset for the detection of the tools in the wild.
1 paper · 0 benchmarks
Minor Irrigation Structures Check-Dam Dataset is a public dataset annotated by domain experts using images from Google static map for instance segmentation and object detection tasks.
1 paper · 0 benchmarks
MlGesture is a dataset for hand gesture recognition tasks, recorded in a car with 5 different sensor types at two different viewpoints.
1 paper · 0 benchmarks
Marine Microalgae Detection in Microscopy Images dataset contains a total number of images in the dataset is 937 and all the objects in these images were annotated.
1 paper · 0 benchmarks
MuCeD, a dataset that is carefully curated and validated by expert pathologists from the All India Institute of Medical Science (AIIMS), Delhi, India.
1 paper · 0 benchmarks
NAO (Natural Adversarial Object)
Natural Adversarial Objects (NAO) is a new dataset to evaluate the robustness of object detection models.
1 paper · 1 benchmark
NII-CU MAPD (NII-CU Multispectral Aerial Person Detection Dataset)
The National Institute of Informatics - Chiba University (NII-CU) Multispectral Aerial Person Detection Dataset consists of 5,880 pairs of aligned RGB+FIR (Far infrared) images captured from a drone flying at heights between 20 and 50…
1 paper · 2 benchmarks
An object-centric version of Stylized COCO to benchmark texture bias and out-of-distribution robustness of vision models.
1 paper · 0 benchmarks
Object detection dataset featuring people walking on grass captured aboard a UAV.
1 paper · 0 benchmarks
PoTATO is a dataset designed to enhance the detection of floating plastic waste in aquatic environments by leveraging polarimetric imaging.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RASMD (RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions)
Current autonomous driving algorithms heavily rely on the visible spectrum, which is prone to performance degradation in adverse conditions like fog, rain, snow, glare, and high contrast.
1 paper · 0 benchmarks
A multimodal dataset of radio galaxies and their corresponding infrared hosts.
1 paper · 0 benchmarks
A dataset to encourage the community to adapt oriented bounding box (OBB) detectors for more complex environments.
1 paper · 0 benchmarks
S-ODv2 (SeaDronesSee-Object Detection v2)
SeaDronesSee-Object Detection v2 (S-ODv2) dataset contains 14,227 RGB images (training: 8,930; validation: 1,547; testing: 3,750).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
This work contributes a large, complex, and realistic high-quality safety clothing and helmet detection (SFCHD) dataset.
1 paper · 1 benchmark
The dataset SFU-HW-Objects-v1 contains bounding boxes and object class labels for High Efficiency Video Coding (HEVC) v1 Common Test Conditions (CTC) video sequences.
1 paper · 0 benchmarks
SIDOD is a new, publicly-available image dataset generated by the NVIDIA Deep Learning Data Synthesizer intended for use in object detection, pose estimation, and tracking applications.
1 paper · 0 benchmarks
SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
1 paper · 0 benchmarks
The Synthetic Signature Bankcheck Images (SSBI) Dataset is the first publicly available dataset of bank check images with annotations for detecting handwritten components, including names, amounts, dates, and signatures.
1 paper · 0 benchmarks
This is a dataset to benchmark real-time embedded object detection models for RoboCup SSL (Small Size League).
1 paper · 0 benchmarks
STN PLAD (STN Power Line Assets Dataset)
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components.
1 paper · 1 benchmark
SemanticSugarBeets, a novel and high-quality dataset containing 953 monocular RGB images and 2920 annotations of sugar beets, enables a wide range of learning tasks including object detection, semantic segmentation, instance segmentation…
1 paper · 0 benchmarks
Songdo Vision (Songdo Vision: Vehicle Annotations from High-Altitude BeV Drone Imagery in a Smart City)
The Songdo Vision dataset provides high-resolution (4K, 3840×2160 pixels) RGB images annotated with categorized axis-aligned bounding boxes (BBs) for vehicle detection from a high-altitude bird’s-eye view (BeV) perspective.
1 paper · 1 benchmark
Synthetic soccer players rendered on top of real world stadium images in 4K covering half a pitch each.
1 paper · 1 benchmark
A real-world image dataset that contains more than 900 images generated from 26 street cameras and 7 object categories annotated with detailed bounding box.
1 paper · 0 benchmarks
SunspotsYoloDataset is a set of 1690+380+128 high-resolution RGB astronomical images captured with smart telescopes with specific solar filters and annotated with the positions of sunspots that are effectively in the images.
1 paper · 0 benchmarks
TAMPAR is a real-world dataset of parcel photos for tampering detection with annotations in COCO format.
1 paper · 0 benchmarks
Please refer this paper @article{sarker2024tea, author = {Sarker, Swapnil Sharma and Islam, Ashiqul and Talukder Raktim, Raufun and Roshni, Sanjana and Joy, Sajib Kumar Saha and Shah, Faisal}, title = {Real-Time Tea Leaf Disease Detection…
1 paper · 0 benchmarks
TiROD (Tiny Robotics Object Detection)
Dataset to benchmark Continual Learning for Object Detection in a Tiny Robotics settings.
1 paper · 1 benchmark
The TimberVision dataset consists of more than 2k annotated RGB images and contains a total of 51k trunk components including cut and lateral surfaces, thereby surpassing any existing dataset in this domain in terms of both quantity and…
1 paper · 0 benchmarks
Mapping urban large-area advertising structures using drone imagery and deep learning-based spatial data analysis.
1 paper · 1 benchmark
The UAVVaste dataset consists to date of 772 images and 3716 annotations.
1 paper · 1 benchmark
UDA-CH (Unsupervised Domain Adaptation on Cultural Heritage)
UDA-CH contains 16 objects that cover a variety of artworks which can be found in a museum like sculptures, paintings and books.
1 paper · 1 benchmark
USB (Universal-Scale Object Detection Benchmark)
The Universal-Scale object detection Benchmark (USB) is a benchmark for object detection that has variations in object scales and image domains by incorporating COCO with the recently proposed Waymo Open Dataset and Manga109-s dataset.
1 paper · 1 benchmark
VETRA is a dataset for vehicle tracking in aerial image sequences and presents unique challenges such as low frame rates, small and fast-moving objects, as well as high camera movement.
1 paper · 0 benchmarks
VME & CDSI (Vehicles in the Middle East (VME) & Car Detection in Satellite Imagery (CDSI) datasets)
Vehicles in the Middle East (VME) dataset, designed explicitly for vehicle detection in high-resolution satellite images from Middle Eastern countries.
1 paper · 1 benchmark
VQA-OV (Visual Quality Assessment of Omnidirectional Video)
Collects 60 reference sequences and 540 impaired sequences.
1 paper · 0 benchmarks
Dataset for testing the ability of Vision Language Models (LVM) to recognize and match 3D objects of the exact same 3D shapes but with different orientation/materials/textures/ environments and light conditions.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.