Home › Datasets › task › Instance Segmentation

Instance Segmentation datasets

archive 2025-07-28

111 datasets carry the task tag "Instance Segmentation" (the task itself: Instance Segmentation), ordered by the archive's paper count. Page 1 of 3: 48 shown of 111. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Instance Segmentation datasets 1–48 of 111

The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
The ADE20K semantic segmentation dataset contains more than 20K scene-centric images exhaustively annotated with pixel-level objects and object parts labels.
1,213 papers · 32 benchmarks
NYUv2 (NYU-Depth V2)
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
KITTI-360 is a large-scale dataset that contains rich sensory information and full annotations.
246 papers · 7 benchmarks
YouTubeVIS is a new dataset tailored for tasks like simultaneous detection, segmentation and tracking of object instances in videos and is collected based on the current largest video object segmentation dataset YouTubeVOS.
163 papers · 2 benchmarks
Objects365 is a large-scale object detection dataset, Objects365, which has 365 object categories over 600K training images.
161 papers · 2 benchmarks
PartNet is a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information.
156 papers · 3 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
iSAID contains 655,451 object instances for 15 categories across 2,806 high-resolution images.
81 papers · 4 benchmarks
3D-FUTURE (3D FUrniture shape with TextURE) is a 3D dataset that contains 20,240 photo-realistic synthetic images captured in 5,000 diverse scenes, and 9,992 involved unique industrial 3D CAD shapes of furniture with high-resolution…
48 papers · 0 benchmarks
WildDash is a benchmark evaluation method is presented that uses the meta-information to calculate the robustness of a given algorithm with respect to the individual hazards.
47 papers · 2 benchmarks
Synscapes is a synthetic dataset for street scene parsing created using photorealistic rendering techniques, and show state-of-the-art results for training and validation as well as new types of analysis.
46 papers · 1 benchmark
The SIXray dataset is constructed by the Pattern Recognition and Intelligent System Development Laboratory, University of Chinese Academy of Sciences.
39 papers · 1 benchmark
Our project (STPLS3D) aims to provide a large-scale aerial photogrammetry dataset with synthetic and real annotated 3D point clouds for semantic and instance segmentation tasks.
36 papers · 3 benchmarks
OCID (Object Clutter Indoor Dataset)
Developing robot perception systems for handling objects in the real-world requires computer vision algorithms to be carefully scrutinized with respect to the expected operating domain.
29 papers · 1 benchmark
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
ScribbleSup (PASCAL-Scribble Dataset)
The PASCAL-Scribble Dataset is an extension of the PASCAL dataset with scribble annotations for semantic segmentation.
26 papers · 0 benchmarks
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
Hi4D contains 4D textured scans of 20 subject pairs, 100 sequences, and a total of more than 11K frames.
19 papers · 0 benchmarks
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
CryoNuSeg is a fully annotated FS-derived cryosectioned and H&E-stained nuclei instance segmentation dataset.
16 papers · 0 benchmarks
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
Fashionpedia consists of two parts: (1) an ontology built by fashion experts containing 27 main apparel categories, 19 apparel parts, 294 fine-grained attributes and their relationships; (2) a dataset with everyday and celebrity event…
14 papers · 0 benchmarks
TTPLA (Transmission Towers and Power Lines (TTPLA))
TTPLA is a public dataset which is a collection of aerial images on Transmission Towers (TTs) and Power Lines (PLs).
13 papers · 0 benchmarks
A Multi-Task 4D Radar-Camera Fusion Dataset for Autonomous Driving on Water Surfaces description of the dataset WaterScenes, the first multi-task 4D radar-camera fusion dataset on water surfaces, which offers data from multiple sensors,…
13 papers · 2 benchmarks
The TrashCan dataset is an instance-segmentation dataset of underwater trash.
12 papers · 0 benchmarks
TACO is a growing image dataset of waste in the wild.
11 papers · 0 benchmarks
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
SpaceNet 2 (SpaceNet 2: Building Detection v2)
SpaceNet 2: Building Detection v2 - is a dataset for building footprint detection in geographically diverse settings from very high resolution satellite images.
10 papers · 1 benchmark
The AIRS (Aerial Imagery for Roof Segmentation) dataset provides a wide coverage of aerial imagery with 7.5 cm resolution and contains over 220,000 buildings.
9 papers · 1 benchmark
PhenoBench (PhenoBench — A Large Dataset and Benchmarks for Semantic Image Interpretation in the Agricultural Domain)
The PhenoBench dataset contains multiple image segmentation challenges from the agricultural domain.
9 papers · 0 benchmarks
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
ROBUST-MIS (Robust Medical Instrument Segmentation Challenge 2019)
The ROBUST-MIS dataset was made available to support the Robust Medical Instrument Segmentation (ROBUST-MIS) Challenge 2019, part of the Endoscopic Vision Challenge associated with MICCAI.
8 papers · 1 benchmark
TikTok Dataset (Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos)
We learn high fidelity human depths by leveraging a collection of social media dance videos scraped from the TikTok mobile social networking application.
8 papers · 0 benchmarks
NDD20 (Northumberland Dolphin Dataset 2020)
Northumberland Dolphin Dataset 2020 (NDD20) is a challenging image dataset annotated for both coarse and fine-grained instance segmentation and categorisation.
7 papers · 0 benchmarks
Augments the KITTI with more instance pixel-level annotation for 8 categories.
6 papers · 1 benchmark
LIS (low-light instance segmentation)
To reveal and systematically investigate the effectiveness of the proposed method in the real world, a real low-light image dataset for instance segmentation is necessary and urgently needed.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
ShipSG (Ship Segmentation and Georeferencing Dataset)
The ShipSG dataset is the first public dataset of its kind for ship segmentation and georeferencing.
6 papers · 0 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
Consists of user-generated aerial videos from social media with annotations of instance-level building damage masks.
5 papers · 0 benchmarks
NERDS 360 (NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes)
We present a large-scale dataset for 3D urban scene understanding.
5 papers · 0 benchmarks
OoDIS (Anomaly Instance Segmentation Benchmark)
OoDIS is a benchmark dataset for anomaly instance segmentation, crucial for autonomous vehicle safety.
5 papers · 2 benchmarks
SpaceNet 1 (SpaceNet 1: Building Detection v1)
SpaceNet 1: Building Detection v1 is a dataset for building footprint detection.
5 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.