Home › Datasets › task › Semantic Segmentation
Semantic Segmentation datasets
archive 2025-07-28
347 datasets carry the task tag "Semantic Segmentation" (the task itself: Semantic Segmentation), ordered by the archive's paper count. Page 1 of 8: 48 shown of 347. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Semantic Segmentation datasets 1–48 of 347
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
ShapeNet is a large scale repository for 3D CAD models developed by researchers from Stanford University, Princeton University and the Toyota Technological Institute at Chicago, USA.
1,947 papers · 13 benchmarks
ScanNet is an instance-level indoor RGB-D dataset that includes both 2D and 3D data.
1,595 papers · 21 benchmarks
The ADE20K semantic segmentation dataset contains more than 20K scene-centric images exhaustively annotated with pixel-level objects and object parts labels.
1,213 papers · 32 benchmarks
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
Eurosat is a dataset and deep learning benchmark for land use and land cover classification.
687 papers · 8 benchmarks
SYNTHIA (SYNTHetic Collection of Imagery and Annotations)
The SYNTHIA dataset is a synthetic dataset that consists of 9400 multi-viewpoint photo-realistic frames rendered from a virtual city and comes with pixel-level semantic annotations for 13 classes.
538 papers · 10 benchmarks
S3DIS (Stanford 3D Indoor Scene Dataset (S3DIS))
The Stanford 3D Indoor Scene Dataset (S3DIS) dataset contains 6 large-scale indoor areas with 271 rooms.
488 papers · 9 benchmarks
The SUN RGBD dataset contains 10335 real RGB-D images of room scenes.
477 papers · 11 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images.
439 papers · 11 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
GTA5 (Grand Theft Auto 5)
The GTA5 dataset contains 24966 synthetic images with pixel level semantic annotation.
412 papers · 7 benchmarks
The Common Objects in COntext-stuff (COCO-stuff) dataset is a dataset for scene understanding tasks like semantic segmentation, object detection and image captioning.
338 papers · 17 benchmarks
The PASCAL Context dataset is an extension of the PASCAL VOC 2010 detection challenge, and it contains pixel-wise labels for all training images.
323 papers · 6 benchmarks
DAVIS17 is a dataset for video object segmentation.
308 papers · 12 benchmarks
KITTI-360 is a large-scale dataset that contains rich sensory information and full annotations.
246 papers · 7 benchmarks
CamVid (Cambridge-driving Labeled Video Database)
CamVid (Cambridge-driving Labeled Video Database) is a road/driving scene understanding database which was originally captured as five video sequences with a 960×720 resolution camera mounted on the dashboard of a car.
227 papers · 4 benchmarks
VisDA-2017 is a simulation-to-real dataset for domain adaptation with over 280,000 images across 12 categories in the training, validation and testing domains.
223 papers · 6 benchmarks
HAM10000 is a dataset of 10000 training images for detecting pigmented skin lesions.
209 papers · 2 benchmarks
The HELEN dataset is composed of 2330 face images of 400×400 pixels with labeled facial components generated through manually-annotated contours along eyes, eyebrows, nose, lips and jawline.
201 papers · 1 benchmark
Kvasir-SEG is an open-access dataset of gastrointestinal polyp images and corresponding segmentation masks, manually annotated by a medical doctor and then verified by an experienced gastroenterologist.
201 papers · 2 benchmarks
PASCAL VOC (PASCAL Visual Object Classes Challenge)
The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa,…
198 papers · 18 benchmarks
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
LabelMe database is a large collection of images with ground truth labels for object detection and recognition.
178 papers · 1 benchmark
PASCAL-5i is a dataset used to evaluate few-shot segmentation.
177 papers · 1 benchmark
YouTubeVIS is a new dataset tailored for tasks like simultaneous detection, segmentation and tracking of object instances in videos and is collected based on the current largest video object segmentation dataset YouTubeVOS.
163 papers · 2 benchmarks
Objects365 is a large-scale object detection dataset, Objects365, which has 365 object categories over 600K training images.
161 papers · 2 benchmarks
PartNet is a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information.
156 papers · 3 benchmarks
The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations.
147 papers · 6 benchmarks
STARE (Structured Analysis of the Retina)
The STARE (Structured Analysis of the Retina) dataset is a dataset for retinal vessel segmentation.
146 papers · 6 benchmarks
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
The Make3D dataset is a monocular Depth Estimation dataset that contains 400 single training RGB and depth map pairs, and 134 test samples.
129 papers · 1 benchmark
PASCAL VOC 2007 is a dataset for image recognition.
126 papers · 13 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
116 papers · 4 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
SegTrack v2 is a video segmentation dataset with full pixel-level annotations on multiple objects at each frame within each video.
107 papers · 5 benchmarks
AVE (Audio-Visual Event Localization)
To investigate three temporal localization tasks: supervised and weakly-supervised audio-visual event localization, and cross-modality localization.
103 papers · 0 benchmarks
Mapillary Vistas Dataset is a diverse street-level imagery dataset with pixel‑accurate and instance‑specific human annotations for understanding street scenes around the world.
101 papers · 3 benchmarks
IDD (Indian Driving Dataset)
IDD is a dataset for road scene understanding in unstructured environments used for semantic segmentation and object detection for autonomous driving.
98 papers · 1 benchmark
The Medical Segmentation Decathlon is a collection of medical image segmentation datasets.
97 papers · 1 benchmark
BigEarthNet consists of 590,326 Sentinel-2 image patches, each of which is a section of i) 120x120 pixels for 10m bands; ii) 60x60 pixels for 20m bands; and iii) 20x20 pixels for 60m bands.
85 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.