Datasets › COCO-Stuff

COCO-Stuff (Common Objects in COntext-stuff)

Introduced by Holger Caesar et al. in COCO-Stuff: Thing and Stuff Classes in Context1 Jan 2018 archive 2025-07-28

The Common Objects in COntext-stuff (COCO-stuff) dataset is a dataset for scene understanding tasks like semantic segmentation, object detection and image captioning. It is constructed by annotating the original COCO dataset, which originally annotated things while neglecting stuff annotations. There are 164k images in COCO-stuff dataset that span over 172 categories including 80 things, 91 stuff, and 1 unlabeled class.

Source: Image Colorization: A Survey and Dataset Image Source: https://github.com/nightrome/cocostuff

Benchmarks archive 2025-07-28

All 17 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Unsupervised Semantic Segmentation COCO-Stuff-27 DynaSeg - FSF (ResNet-18 FPN) Clustering [mIoU] 54.1 DynaSeg: A Deep Dynamic Fusion Method for Unsupervised... ryersonmultimedialab/dynaseg 29 Compare
Semantic Segmentation COCO-Stuff test VPNeXt mIoU 53.7 VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer — 21 Compare
Image-to-Image Translation COCO-Stuff Labels-to-Photos DP-SIMS (ConvNext-XL) FID 13.3 Unlocking Pre-trained Image Backbones for Semantic Image... — 15 Compare
Zero-Shot Semantic Segmentation COCO-Stuff OTSeg+ Transductive Setting hIoU 49.8 OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot... cubeyoung/OTSeg 15 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training COCO-Stuff-171 CorrCLIP mIoU 34.0 CorrCLIP: Reconstructing Correlations in CLIP with... zdk258/CorrCLIP 12 Compare
Open Vocabulary Semantic Segmentation COCO-Stuff-171 POMP HIoU 39.1 Prompt Pre-Training with Twenty-Thousand Classes for... amazon-science/prompt-pretraining 7 Compare
Layout-to-Image Generation COCO-Stuff 64x64 OC-GAN FID 29.57 Object-Centric Image Generation from Layouts — 5 Compare
Layout-to-Image Generation COCO-Stuff 128x128 CAL2IM FID 22.32 Context-Aware Layout to Image Generation with Enhanced... wtliao/layout2img 5 Compare
Unsupervised Semantic Segmentation COCO-Stuff-171 CAUSE-TR (ViT-S/8) mIoU 15.2 Causal Unsupervised Semantic Segmentation ByungKwanLee/Causal-Unsupervised-Segmentation 4 Compare
Unsupervised Semantic Segmentation COCO-Stuff-81 CAUSE-TR (ViT-S/8) mIoU 21.2 Causal Unsupervised Semantic Segmentation ByungKwanLee/Causal-Unsupervised-Segmentation 4 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training COCO-Stuff-27 ReCo+ mIoU 32.6 ReCo: Retrieve and Co-segment for Zero-shot Transfer NoelShin/reco +1 4 Compare
Sketch-to-Image Translation COCO-Stuff PITI FID 18.5 Pretraining is All You Need for Image-to-Image Translation PITI-Synthesis/PITI +1 3 Compare
Unsupervised Semantic Segmentation COCO-Stuff-15 InfoSeg Pixel Accuracy 38.8 InfoSeg: Unsupervised Semantic Image Segmentation with... — 3 Compare
Real-Time Semantic Segmentation COCO-Stuff BiSeNet V2-Large Frame (fps) 42.5(1080Ti) BiSeNet V2: Bilateral Network with Guided Aggregation... PaddlePaddle/PaddleSeg +6 2 Compare
Semantic Segmentation COCO-Stuff Deeplab v2 F.W. IU 47.6 COCO-Stuff: Thing and Stuff Classes in Context kazuto1011/deeplab-pytorch +9 1 Compare
Semantic Segmentation COCO-Stuff-27 DiffSeg (512) Pixel Accuracy 72.5 Diffuse, Attend, and Segment: Unsupervised Zero-Shot... google/diffseg 1 Compare
Semantic Segmentation COCO-Stuff full SegFormer-B5 (Single Scale) Mean IoU (class) 46.7 SegFormer: Simple and Efficient Design for Semantic... huggingface/transformers +27 1 Compare

Papers archive 2025-07-28

30 shown of 92 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 338. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models 1 1 29 May 2025 not harvested
The Missing Point in Vision Transformers for Universal Image Segmentation 1 1 26 May 2025 not harvested
Hierarchical Context Learning of object components for unsupervised semantic segmentation 1 2 29 Apr 2025 not harvested
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer 0 1 23 Feb 2025 not harvested
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 1 1 2 Dec 2024 not harvested
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 1 1 15 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 1 1 14 Nov 2024 not harvested
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation 1 1 9 Aug 2024 ran 3 of 9 samples (6 unverified; 9 pointer-only for licence)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation 1 1 9 Aug 2024 not harvested
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut 1 1 5 Jun 2024 ran 1 of 2 samples (1 unverified)
DynaSeg: A Deep Dynamic Fusion Method for Unsupervised Image Segmentation Incorporating Feature Similarity and Spatial Continuity 1 1 9 May 2024 not harvested
Boosting Unsupervised Semantic Segmentation with Principal Mask Proposals 1 2 25 Apr 2024 not harvested
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias 1 4 30 Mar 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation 1 2 21 Mar 2024 ran 9 of 10 samples (1 unverified; 10 pointer-only for licence)
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation 1 1 3 Mar 2024 ran 14 of 19 samples (5 unverified)
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis 1 1 26 Feb 2024 ran 1 of 4 samples (3 unverified)
Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation 1 1 1 Jan 2024 not harvested
Unsupervised Universal Image Segmentation 2 1 28 Dec 2023 ran 13 of 15 samples (2 unverified)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification 1 1 21 Dec 2023 not harvested
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis 0 2 20 Dec 2023 not harvested
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training 1 1 20 Dec 2023 not harvested
Expand-and-Quantize: Unsupervised Semantic Segmentation Using High-Dimensional Space and Product Quantization 0 2 12 Dec 2023 not harvested
Causal Unsupervised Semantic Segmentation 1 5 11 Oct 2023 not harvested
Unsupervised Semantic Segmentation Through Depth-Guided Feature Correlation and Sampling 1 3 21 Sep 2023 not harvested
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion 1 1 23 Aug 2023 ran 3 of 3 samples (0 unverified)
Wavelet-based Unsupervised Label-to-Image Translation 1 1 16 May 2023 not harvested
MVP-SEG: Multi-View Prompt Learning for Open-Vocabulary Semantic Segmentation 0 1 14 Apr 2023 not harvested
A Closer Look at the Explainability of Contrastive Language-Image Pre-training 2 1 12 Apr 2023 not harvested
Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition 1 1 10 Apr 2023 ran 5 of 7 samples (2 unverified)
Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network 1 1 3 Apr 2023 not harvested

The full list of 92 is in the JSON twin.

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Various

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • COCO-Stuff-81
  • COCO-Stuff 256x256
  • COCO-Stuff-171
  • COCO-Stuff-27
  • COCO-Stuff full
  • COCO-Stuff-3
  • COCO-Stuff-15
  • COCO-Stuff test
  • COCO-Stuff Labels-to-Photos
  • COCO-Stuff 64x64
  • COCO-Stuff 128x128
  • COCO-Stuff

12 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections