Datasets › ADE20K

ADE20K

Introduced by Bolei Zhou et al. in Scene Parsing Through ADE20K Dataset1 Jan 2017 archive 2025-07-28

The ADE20K semantic segmentation dataset contains more than 20K scene-centric images exhaustively annotated with pixel-level objects and object parts labels. There are totally 150 semantic categories, which include stuffs like sky, road, grass, and discrete objects like person, car, bed.

Source: Cooperative Image Segmentation and Restoration in Adverse Environmental Conditions Image Source: https://groups.csail.mit.edu/vision/datasets/ADE20K/

Benchmarks archive 2025-07-28

All 32 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Semantic Segmentation ADE20K ViT-P (InternImage-H) Validation mIoU 63.6 The Missing Point in Vision Transformers for Universal... sajjad-sh33/vit-p 235 Compare
Semantic Segmentation ADE20K val BEiT-3 mIoU 62.8 Image as a Foreign Language: BEiT Pretraining for All... microsoft/unilm +1 95 Compare
Panoptic Segmentation ADE20K val OneFormer (InternImage-H, emb_dim=256, single-scale, 896x896) PQ 54.5 OneFormer: One Transformer to Rule Universal Image Segmentation huggingface/transformers +3 25 Compare
Open Vocabulary Semantic Segmentation ADE20K-150 Mask-Adapter mIoU 38.2 Mask-Adapter: The Devil is in the Masks for... hustvl/maskadapter 23 Compare
Open Vocabulary Semantic Segmentation ADE20K-847 UMG-CLIP-E/14 mIoU 17.3 UMG-CLIP: A Unified Multi-Granularity Vision Generalist... lygsbw/umg-clip 19 Compare
Image-to-Image Translation ADE20K Labels-to-Photos DP-SIMS (ConvNext-L) mIoU 54.3 Unlocking Pre-trained Image Backbones for Semantic Image... — 16 Compare
Instance Segmentation ADE20K val OneFormer (InternImage-H, emb_dim=1024, single-scale, 896x896, COCO-Pretrained) AP 44.2 OneFormer: One Transformer to Rule Universal Image Segmentation huggingface/transformers +3 14 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training ADE20K CorrCLIP Mean IoU (val) 30.7 CorrCLIP: Reconstructing Correlations in CLIP with... zdk258/CorrCLIP 13 Compare
Open Vocabulary Panoptic Segmentation ADE20K UMG-CLIP-E/14 PQ 31.6 UMG-CLIP: A Unified Multi-Granularity Vision Generalist... lygsbw/umg-clip 10 Compare
Overlapped 100-5 ADE20K MBS mIoU 42.8 Mitigating Background Shift in Class-Incremental... roadonep/eccv2024_mbs 8 Compare
Image-to-Image Translation ADE20K-Outdoor Labels-to-Photos DP-GAN mIoU 40.4 Dual Pyramid Generative Adversarial Networks for... sj-li/dp_gan 7 Compare
Overlapped 100-50 ADE20K MBS mIoU 45.7 Mitigating Background Shift in Class-Incremental... roadonep/eccv2024_mbs 7 Compare
Overlapped 50-50 ADE20K MBS mIoU 45.4 Mitigating Background Shift in Class-Incremental... roadonep/eccv2024_mbs 7 Compare
Overlapped 100-10 ADE20K MBS Mean IoU (test) 44.5 Mitigating Background Shift in Class-Incremental... roadonep/eccv2024_mbs 6 Compare
Semi-Supervised Semantic Segmentation ADE20K 1/32 labeled UniMatch V2 Validation mIoU 45.0 UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 5 Compare
Semi-Supervised Semantic Segmentation ADE20K 1/16 labeled UniMatch V2 Validation mIoU 46.7 UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 5 Compare
Sound Prompted Semantic Segmentation ADE20K DenseAV mAP 32.7 Separating the "Chirp" from the "Chat": Self-supervised... mhamilton723/DenseAV 4 Compare
Speech Prompted Semantic Segmentation ADE20K DenseAV mAP 48.7 Separating the "Chirp" from the "Chat": Self-supervised... mhamilton723/DenseAV 4 Compare
Continual Semantic Segmentation ADE20K LGKD mIoU 37.5 Label-Guided Knowledge Distillation for Continual... ze-yang/lgkd 2 Compare
Face Detection ADE20K CASSOD mIoU 42.86 CASSOD-Net: Cascaded and Separable Structures of Dilated... — 1 Compare
Overlapped 25-25 ADE20K SATS-M Mean IoU (test) 32.56 SATS: Self-Attention Transfer for Continual Semantic Segmentation QIU023/SATS_Continual_Semantic_Seg 1 Compare
Panoptic Segmentation ADE20K MasQCLIP PQ 23.3 MasQCLIP for Open-Vocabulary Universal Image Segmentation mlpc-ucsd/MasQCLIP 1 Compare
Pose Transfer ADE20K SCAM FID 27.5 SCAM! Transferring humans between images with Semantic... nicolas-dufour/SCAM 1 Compare
Reconstruction ADE20K SCAM PSNR 20 SCAM! Transferring humans between images with Semantic... nicolas-dufour/SCAM 1 Compare
Scene Recognition ADE20K Semantic-Aware Scene Recogniton (ResNet-18) Top 1 Accuracy 62.55 Semantic-Aware Scene Recognition vpulab/Semantic-Aware-Scene-Recognition 1 Compare
Scene Understanding ADE20K val CPN(ResNet-101) Mean IoU 46.3 Context Prior for Scene Segmentation ycszen/ContextPrior +1 1 Compare
Semi-Supervised Instance Segmentation ADE20K CAST AP 16.7 CAST: Contrastive Adaptation and Distillation for... — 1 Compare
Weakly-Supervised Semantic Segmentation ADE20K val DHR (Swin-L, Mask2Former) mIoU 32.9 DHR: Dual Features-Driven Hierarchical Rebalancing in... shjo-april/DHR 1 Compare
Zero-Shot Semantic Segmentation ADE20K-847 MAFT unseen mIoU 8.7 — — 1 Compare
Open-Vocabulary Semantic Segmentation ADE20K-150 no rows — — 0 Compare
Open Vocabulary Semantic Segmentation ADE20K-150 no rows — — 0 Compare
Semantic Segmentation ADE20K-150 no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 198 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,213. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models 1 1 29 May 2025 not harvested
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation 0 1 28 May 2025 not harvested
The Missing Point in Vision Transformers for Universal Image Segmentation 1 7 26 May 2025 not harvested
Your ViT is Secretly an Image Segmentation Model 1 3 24 Mar 2025 ran 1 of 2 samples (1 unverified)
MaskCLIP++: A Mask-Based CLIP Fine-tuning Framework for Open-Vocabulary Image Segmentation 1 2 16 Dec 2024 not harvested
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation 1 2 5 Dec 2024 not harvested
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 1 1 2 Dec 2024 not harvested
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 1 1 15 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 1 1 14 Nov 2024 not harvested
Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression Perspective 1 2 5 Nov 2024 not harvested
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation 1 2 14 Oct 2024 ran 2 of 5 samples (3 unverified)
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention 1 1 11 Oct 2024 not harvested
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation 1 1 9 Aug 2024 ran 3 of 9 samples (6 unverified; 9 pointer-only for licence)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation 1 1 9 Aug 2024 not harvested
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation 1 3 1 Aug 2024 not harvested
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders 1 1 17 Jul 2024 not harvested
Mitigating Background Shift in Class-Incremental Semantic Segmentation 1 4 16 Jul 2024 not harvested
Open-Vocabulary Semantic Segmentation with Image Embedding Balancing 1 2 14 Jun 2024 ran 3 of 12 samples (9 unverified)
Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language 1 2 9 Jun 2024 ran 2 of 3 samples (1 unverified)
Parameter-Inverted Image Pyramid Networks 1 1 6 Jun 2024 not harvested
OpenDAS: Open-Vocabulary Domain Adaptation for 2D and 3D Segmentation 0 1 30 May 2024 not harvested
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias 1 4 30 Mar 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
DHR: Dual Features-Driven Hierarchical Rebalancing in Inter- and Intra-Class Regions for Weakly-Supervised Semantic Segmentation 1 1 30 Mar 2024 ran 9 of 9 samples (0 unverified; 9 pointer-only for licence)
PosSAM: Panoptic Open-vocabulary Segment Anything 1 2 14 Mar 2024 not harvested
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions 2 1 13 Mar 2024 not harvested
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis 1 1 26 Feb 2024 ran 1 of 4 samples (3 unverified)
SERNet-Former: Semantic Segmentation by Efficient Residual Network with Attention-Boosting Gates and Attention-Fusion Networks 2 2 28 Jan 2024 not harvested
UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding 1 6 12 Jan 2024 not harvested
Harnessing Diffusion Models for Visual Perception with Meta Prompts 1 1 22 Dec 2023 ran 4 of 6 samples (2 unverified)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification 1 1 21 Dec 2023 not harvested

The full list of 198 is in the JSON twin.

Dataset loaders archive 2025-07-28

14 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial)

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • ADE20K-847
  • ADE20K-150
  • ADE20K 1/32 labeled
  • ADE20K 1/16 labeled
  • ADE20K-Outdoor Labels-to-Photos
  • ADE20K val
  • ADE20K Labels-to-Photos
  • ADE20K

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections