Datasets › ImageNet-S

ImageNet-S (ImageNet Semantic Segmentation)

Introduced by ShangHua Gao et al. in Large-scale Unsupervised Semantic Segmentation6 Jun 2021 archive 2025-07-28

Powered by the ImageNet dataset, unsupervised learning on large-scale data has made significant advances for classification tasks. There are two major challenges to allowing such an attractive learning modality for segmentation tasks: i) a large-scale benchmark for assessing algorithms is missing; ii) unsupervised shape representation learning is difficult. We propose a new problem of large-scale unsupervised semantic segmentation (LUSS) with a newly created benchmark dataset to track the research progress. Based on the ImageNet dataset, we propose the ImageNet-S dataset with 1.2 million training images and 50k high-quality semantic segmentation annotations for evaluation. Our benchmark has a high data diversity and a clear task objective. We also present a simple yet effective baseline method that works surprisingly well for LUSS. In addition, we benchmark related un/weakly/fully supervised methods accordingly, identifying the challenges and possible directions of LUSS.

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 43. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MMRL: Multi-Modal Representation Learning for Vision-Language Models 1 1 11 Mar 2025 ran 0 of 1 samples (1 unverified)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling 2 1 27 Aug 2024 not harvested
Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language Models 2 1 11 Dec 2023 ran 3 of 7 samples (4 unverified)
Self-regulating Prompts: Foundational Model Adaptation without Forgetting 2 1 13 Jul 2023 ran 7 of 21 samples (14 unverified)
Consistency-guided Prompt Learning for Vision-Language Models 2 1 1 Jun 2023 ran 3 of 7 samples (4 unverified)
Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition 1 1 10 Apr 2023 ran 5 of 7 samples (2 unverified)
Towards Sustainable Self-supervised Learning 1 4 20 Oct 2022 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
MaPLe: Multi-modal Prompt Learning 3 1 6 Oct 2022 ran 4 of 6 samples (2 unverified)
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 1 14 Sep 2022 ran 2 of 4 samples (2 unverified)
RF-Next: Efficient Receptive Field Search for Convolutional Neural Networks 2 3 14 Jun 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
SERE: Exploring Feature Self-relation for Self-supervised Transformer 1 6 10 Jun 2022 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
Conditional Prompt Learning for Vision-Language Models 12 1 10 Mar 2022 ran 4 of 6 samples (2 unverified)
A ConvNet for the 2020s 54 1 10 Jan 2022 ran 54 of 80 samples (26 unverified; 11 pointer-only for licence)
Masked Autoencoders Are Scalable Vision Learners 58 4 11 Nov 2021 ran 71 of 137 samples (66 unverified; 73 pointer-only for licence)
Large-scale Unsupervised Semantic Segmentation 3 6 6 Jun 2021 not harvested
PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering 2 1 30 Mar 2021 ran 2 of 4 samples (2 unverified; 1 pointer-only for licence)
Learning Transferable Visual Models From Natural Language Supervision 82 1 26 Feb 2021 ran 16 of 20 samples (4 unverified; 16 pointer-only for licence)
Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals 2 1 11 Feb 2021 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Deep Clustering for Unsupervised Learning of Visual Features 9 1 15 Jul 2018 ran 5 of 7 samples (2 unverified; 4 pointer-only for licence)

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ImageNet-S
  • ImageNet-S-300
  • ImageNet-S-50

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections