Browse State-of-the-Art › Instance Segmentation

Instance Segmentation

1,158 papers with code · 36 benchmarks · 111 datasets archive 2025-07-28

Computer Vision

Instance Segmentation is a computer vision task that involves identifying and separating individual objects within an image, including detecting the boundaries of each object and assigning a unique label to each object. The goal of instance segmentation is to produce a pixel-wise segmentation map of the image, where each pixel is assigned to a specific object instance.

Image Credit: Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers, CVPR'21

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

36 leaderboard tables shown for this task, 36 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 36 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
COCO test-dev (112 rows) Co-DETR DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
COCO minival (93 rows) Co-DETR DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
LVIS v1.0 val (25 rows) Co-DETR (single-scale) DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
Cityscapes val (17 rows) ViT-P (OneFormer, ConvNeXt-L, single-scale, 512x1024, Mapillary Vistas-pretrained) The Missing Point in Vision Transformers for Universal Image Segmentation code — Compare
ADE20K val (14 rows) OneFormer (InternImage-H, emb_dim=1024, single-scale, 896x896, COCO-Pretrained) OneFormer: One Transformer to Rule Universal Image Segmentation code Syntology ran 0 of 5 samples · 5 unverified Compare
Cityscapes test (11 rows) Deep Watershed Transform Deep Watershed Transform for Instance Segmentation code Syntology ran 0 of 8 samples · 8 unverified Compare
ARMBench (7 rows) RISE (VIT-B) Robot Instance Segmentation with Few Annotations for Grasping code — Compare
Occluded COCO (6 rows) Swin-B + Cascade Mask R-CNN (tri-layer modelling) A Tri-Layer Plugin to Improve Occluded Detection code — Compare
Separated COCO (6 rows) Swin-B + Cascade Mask R-CNN (tri-layer modelling) A Tri-Layer Plugin to Improve Occluded Detection code — Compare
iSAID (5 rows) PANet++ iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images code Syntology ran 1 of 1 samples · 0 unverified Compare
TBBR (5 rows) Swin-T (ImageNet-1k pretrain) Deep learning approaches to building rooftop thermal bridge... code — Compare
BDD100K val (4 rows) Mask Transfiner Mask Transfiner for High-Quality Instance Segmentation code Syntology ran 1 of 1 samples · 0 unverified Compare
COCO 2017 val (4 rows) SparK (ConvNeXt V1-B Mask R-CNN) Designing BERT for Convolutional Networks: Sparse and Hierarchical... code Syntology ran 5 of 14 samples · 9 unverified Compare
COCO val (panoptic labels) (4 rows) OneFormer (InternImage-H, emb_dim=1024, single-scale) OneFormer: One Transformer to Rule Universal Image Segmentation code Syntology ran 0 of 5 samples · 5 unverified Compare
OoDIS (3 rows) UGainS UGainS: Uncertainty Guided Anomaly Instance Segmentation code — Compare
NYUDv2-IS (2 rows) IAM + SOLQ IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks code — Compare
SUN-RGBD-IS (2 rows) IAM + SOLQ IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks code — Compare
UIIS (2 rows) WaterMask RCNN WaterMask: Instance Segmentation for Underwater Imagery code — Compare
Box-IS (1 row) IAM + SOLQ IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks code — Compare
Cityscapes (1 row) CAST CAST: Contrastive Adaptation and Distillation for Semi-Supervised... — — Compare
COCO (1 row) ColorMAE-Green-ViTB-1600 ColorMAE: Exploring data-independent masking strategies in Masked... code — Compare
coco minval (1 row) R3-CNN (ResNet-50-FPN, GC-Net) Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing code — Compare
COCO-N Medium (1 row) Mask R-CNN ResNet-50 FPN Benchmarking Label Noise in Instance Segmentation: Spatial Noise Matters code Syntology ran 2 of 2 samples · 0 unverified Compare
COCO val2017 (1 row) MogaNet-S (256x192) MogaNet: Multi-order Gated Aggregation Network code Syntology ran 12 of 15 samples · 3 unverified Compare
iShape (1 row) ASIS(baseline) iShape: A First Step Towards Irregular Shape Instance Segmentation — — Compare
KINS (1 row) BCNet Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers code Syntology ran 7 of 12 samples · 5 unverified Compare
LDD (1 row) R^3-CNN LDD: A Dataset for Grape Diseases Object Detection and Instance... — — Compare
Leaf Segmentation Challenge (1 row) LeafMask LeafMask: Towards Greater Accuracy on Leaf Segmentation code — Compare
LVIS v1.0 test-dev (1 row) R50-FPN-MaskRCNN-TTA 1st Place Solution of LVIS Challenge 2020: A Good Box is not a... — — Compare
nuScenes (1 row) TraDeS Track to Detect and Segment: An Online Multi-Object Tracker code Syntology ran 4 of 4 samples · 0 unverified Compare
NYU Depth v2 (1 row) SGPN-CNN SGPN: Similarity Group Proposal Network for 3D Point Cloud... code — Compare
PartNet (1 row) PE Point Cloud Instance Segmentation using Probabilistic Embeddings — — Compare
TexBiG 2022 test (1 row) VSR (Vison, Semantics and Relation Model) A Dataset for Analysing Complex Document Layouts in the Digital... code — Compare
TexBiG 2023 test (1 row) DetectoRS + LAEM Drawing the Same Bounding Box Twice? Coping Noisy Annotations in... code — Compare
UAVBillboards (1 row) YOLOv8-X Mapping urban large-area advertising structures using drone... code — Compare
UFBA-425 (1 row) BB-UNet Instance Segmentation and Teeth Classification in Panoramic X-rays code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

111 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 111 until expanded.

Subtasks archive 2025-07-28

17 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 1,158 papers with code (2,262 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 20 Mar 2017 179 repositories listed Syntology ran 42 of 140 samples · 98 unverified · 23 pointer-only (licence)
    Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance.
  • 17 Jun 2019 142 repositories listed Syntology ran 14 of 82 samples · 68 unverified
    In this paper, we introduce the various features of this toolbox.
  • 25 Mar 2021 80 repositories listed Syntology ran 108 of 207 samples · 99 unverified · 43 pointer-only (licence)
    This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision.
  • 4 Apr 2019 48 repositories listed Syntology ran 10 of 21 samples · 11 unverified · 6 pointer-only (licence)
    Then we produce instance masks by linearly combining the prototypes with the mask coefficients.
  • 20 Aug 2019 42 repositories listed Syntology ran 3 of 34 samples · 31 unverified · 16 pointer-only (licence)
    High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection.
  • 25 Feb 2019 39 repositories listed Syntology ran 8 of 25 samples · 17 unverified
    We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel.
  • 1 May 2014 38 repositories listed Syntology ran 1 of 7 samples · 6 unverified
    We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding.
  • 19 Apr 2020 36 repositories listed Syntology ran 8 of 48 samples · 40 unverified · 23 pointer-only (licence)
    It is well known that featuremap attention and multi-path representation are important for visual recognition.
  • 3 Dec 2019 36 repositories listed Syntology ran 11 of 43 samples · 32 unverified
    Then we produce instance masks by linearly combining the prototypes with the mask coefficients.
  • 2 Apr 2019 34 repositories listed Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)
    We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.
  • 21 Nov 2017 32 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
    Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time.
  • 27 Nov 2018 26 repositories listed Syntology ran 2 of 13 samples · 11 unverified · 2 pointer-only (licence)
    The superior performance of Deformable Convolutional Networks arises from its ability to adapt to the geometric variations of objects.
  • 10 Dec 2019 24 repositories listed
    We present a new, embarrassingly simple approach to instance segmentation in images.
  • 18 Nov 2021 23 repositories listed Syntology ran 3 of 30 samples · 27 unverified
    Three main techniques are proposed: 1) a residual-post-norm method combined with cosine attention to improve training stability; 2) A log-spaced continuous position bias method to effectively transfer models pre-trained…
  • 15 Feb 2018 22 repositories listed Syntology ran 3 of 34 samples · 31 unverified · 1 pointer-only (licence)
    By doing so, we ensure a lane fitting which is robust against road plane changes, unlike existing approaches that rely on a fixed, pre-defined transformation.
  • 20 Feb 2022 21 repositories listed Syntology ran 0 of 6 samples · 6 unverified
    In this paper, we propose a novel linear attention named large kernel attention (LKA) to enable self-adaptive and long-range correlations in self-attention while avoiding its shortcomings.
  • 30 Nov 2017 21 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)
    We present a new method for synthesizing high-resolution photo-realistic images from semantic label maps using conditional generative adversarial networks (conditional GANs).
  • 26 Apr 2022 19 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 1 pointer-only (licence)
    Cross-entropy loss and focal loss are the most common choices when training deep neural networks for classification problems.
  • 19 May 2017 19 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 3 pointer-only (licence)
    Numerous deep learning applications benefit from multi-task learning with multiple regression and classification objectives.
  • 23 Mar 2020 18 repositories listed Syntology ran 15 of 38 samples · 23 unverified · 24 pointer-only (licence)
    Importantly, we take one step further by dynamically learning the mask head of the object segmenter such that the mask head is conditioned on the location.
  • 14 Dec 2022 14 repositories listed Syntology ran 3 of 20 samples · 17 unverified
    In this paper, we aim to design an efficient real-time object detector that exceeds the YOLO series and is easily extensible for many object recognition tasks such as instance segmentation and rotated object detection.
  • 17 Dec 2019 14 repositories listed Syntology ran 5 of 18 samples · 13 unverified
    We present a new method for efficient high-quality image segmentation of objects and scenes.
  • 11 Sep 2019 14 repositories listed
    In this paper, we challenge the necessity of such hard/soft sampling methods for training accurate deep object detectors.
  • 4 Dec 2018 14 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 1 pointer-only (licence)
    Dot-product attention has wide applications in computer vision and natural language processing.
  • 27 Jan 2021 13 repositories listed Syntology ran 26 of 49 samples · 23 unverified · 8 pointer-only (licence)
    Finally, we present a simple adaptation of the BoTNet design for image classification, resulting in models that achieve a strong performance of 84.
  • 11 Dec 2019 13 repositories listed Syntology ran 1 of 9 samples · 8 unverified · 2 pointer-only (licence)
    The state-of-the-art models for medical image segmentation are variants of U-Net and fully convolutional networks (FCN).
  • 10 Dec 2019 13 repositories listed
    We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search.
  • 8 Oct 2019 13 repositories listed Syntology ran 1 of 8 samples · 7 unverified
    By dissecting the channel attention module in SENet, we empirically show avoiding dimensionality reduction is important for learning channel attention, and appropriate cross-channel interaction can preserve performance…
  • 17 Jun 2021 12 repositories listed Syntology ran 3 of 14 samples · 11 unverified · 3 pointer-only (licence)
    We propose a "transposed" version of self-attention that operates across feature channels rather than tokens, where the interactions are based on the cross-covariance matrix between keys and queries.
  • 8 Jan 2019 12 repositories listed Syntology ran 7 of 12 samples · 5 unverified
    In this work, we perform a detailed study of this minimally extended version of Mask R-CNN with FPN, which we refer to as Panoptic FPN, and show it is a robust and accurate baseline for both tasks.

Syntology lines on 27 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections