{"url":"/dataset/imagenet-s","name":"ImageNet-S","full_name":"ImageNet Semantic Segmentation","description_markdown":"Powered by the ImageNet dataset, unsupervised learning on large-scale data has made significant advances for classification tasks. There are two major challenges to allowing such an attractive learning modality for segmentation tasks: i) a large-scale benchmark for assessing algorithms is missing; ii) unsupervised shape representation learning is difficult. We propose a new problem of large-scale unsupervised semantic segmentation (LUSS) with a newly created benchmark dataset to track the research progress. Based on the ImageNet dataset, we propose the ImageNet-S dataset with 1.2 million training images and 50k high-quality semantic segmentation annotations for evaluation. Our benchmark has a high data diversity and a clear task objective. We also present a simple yet effective baseline method that works surprisingly well for LUSS. In addition, we benchmark related un/weakly/fully supervised methods accordingly, identifying the challenges and possible directions of LUSS.","description_withheld":null,"homepage":"https://github.com/LUSSeg/ImageNet-S","introduced_date":"2021-06-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/large-scale-unsupervised-semantic","title":"Large-scale Unsupervised Semantic Segmentation","first_author":"ShangHua Gao","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Semantic Segmentation","url":"/task/semantic-segmentation","datasets_with_task":"/datasets/task/semantic-segmentation"},{"name":"Semi-Supervised Semantic Segmentation","url":"/task/semi-supervised-semantic-segmentation","datasets_with_task":"/datasets/task/semi-supervised-semantic-segmentation"},{"name":"Unsupervised Semantic Segmentation","url":"/task/unsupervised-semantic-segmentation","datasets_with_task":"/datasets/task/unsupervised-semantic-segmentation"},{"name":"Zero-Shot Transfer Image Classification","url":"/task/zero-shot-transfer-image-classification","datasets_with_task":"/datasets/task/zero-shot-transfer-image-classification"},{"name":"Prompt Engineering","url":"/task/prompt-engineering","datasets_with_task":"/datasets/task/prompt-engineering"}],"languages":[],"variants":["ImageNet-S","ImageNet-S-300","ImageNet-S-50"],"data_loaders":[{"repo":"https://github.com/LUSSeg/ImageNet-S","url":"https://github.com/LUSSeg/ImageNet-S","frameworks":["pytorch"]},{"repo":"https://github.com/LUSSeg/PASS","url":"https://github.com/LUSSeg/PASS","frameworks":["pytorch"]},{"repo":"https://github.com/LUSSeg/ImageNetSegModel","url":"https://github.com/LUSSeg/ImageNetSegModel","frameworks":["pytorch"]}],"num_papers_in_archive":43,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/semantic-segmentation-on-imagenet-s","task":"Semantic Segmentation","dataset_variant":"ImageNet-S","rows":20,"metrics":["mIoU (val)","mIoU (test)"],"first_row_in_archive_order":{"model":"TEC (ViT-B/16, 224x224, SSL+FT, mmseg)","paper":"/paper/towards-sustainable-self-supervised-learning","metrics":{"mIoU (test)":"62.5","mIoU (val)":"63.2"},"code_links":[{"title":"sail-sg/tec","url":"https://github.com/sail-sg/tec"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/prompt-engineering-on-imagenet-s","task":"Prompt Engineering","dataset_variant":"ImageNet-S","rows":9,"metrics":["Top-1 accuracy %"],"first_row_in_archive_order":{"model":"POMP","paper":"/paper/prompt-pre-training-with-twenty-thousand-1","metrics":{"Top-1 accuracy %":"49.8"},"code_links":[{"title":"amazon-science/prompt-pretraining","url":"https://github.com/amazon-science/prompt-pretraining"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-6","task":"Unsupervised Semantic Segmentation","dataset_variant":"ImageNet-S-50","rows":5,"metrics":["mIoU (test)","mIoU (val)"],"first_row_in_archive_order":{"model":"PASS (+Saliency map)","paper":"/paper/large-scale-unsupervised-semantic","metrics":{"mIoU (test)":"42.3","mIoU (val)":"43.3"},"code_links":[{"title":"LUSSeg/ImageNet-S","url":"https://github.com/LUSSeg/ImageNet-S"},{"title":"LUSSeg/PASS","url":"https://github.com/LUSSeg/PASS"},{"title":"LUSSeg/ImageNetSegModel","url":"https://github.com/LUSSeg/ImageNetSegModel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-4","task":"Unsupervised Semantic Segmentation","dataset_variant":"ImageNet-S","rows":1,"metrics":["mIoU (test)","mIoU (val)"],"first_row_in_archive_order":{"model":"PASS","paper":"/paper/large-scale-unsupervised-semantic","metrics":{"mIoU (test)":"11.0","mIoU (val)":"11.5"},"code_links":[{"title":"LUSSeg/ImageNet-S","url":"https://github.com/LUSSeg/ImageNet-S"},{"title":"LUSSeg/PASS","url":"https://github.com/LUSSeg/PASS"},{"title":"LUSSeg/ImageNetSegModel","url":"https://github.com/LUSSeg/ImageNetSegModel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-5","task":"Unsupervised Semantic Segmentation","dataset_variant":"ImageNet-S-300","rows":1,"metrics":["mIoU (test)","mIoU (val)"],"first_row_in_archive_order":{"model":"PASS","paper":"/paper/large-scale-unsupervised-semantic","metrics":{"mIoU (test)":"18.1","mIoU (val)":"18"},"code_links":[{"title":"LUSSeg/ImageNet-S","url":"https://github.com/LUSSeg/ImageNet-S"},{"title":"LUSSeg/PASS","url":"https://github.com/LUSSeg/PASS"},{"title":"LUSSeg/ImageNetSegModel","url":"https://github.com/LUSSeg/ImageNetSegModel"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/zero-shot-transfer-image-classification-on-9","task":"Zero-Shot Transfer Image Classification","dataset_variant":"ImageNet-S","rows":1,"metrics":["Accuracy (Private)","Top 5 Accuracy"],"first_row_in_archive_order":{"model":"PaLI","paper":"/paper/pali-a-jointly-scaled-multilingual-language","metrics":{"Accuracy (Private)":"63.83","Top 5 Accuracy":"79.3"},"code_links":[{"title":"google-research/big_vision","url":"https://github.com/google-research/big_vision"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/mmrl-multi-modal-representation-learning-for","title":"MMRL: Multi-Modal Representation Learning for Vision-Language Models","date":"2025-03-11","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hpt-hierarchically-prompting-vision-language","title":"HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling","date":"2024-08-27","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/learning-hierarchical-prompt-with-structured","title":"Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language Models","date":"2023-12-11","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":3,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/self-regulating-prompts-foundational-model","title":"Self-regulating Prompts: Foundational Model Adaptation without Forgetting","date":"2023-07-13","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":21,"samples_ran":7,"samples_unverified":14,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/consistency-guided-prompt-learning-for-vision","title":"Consistency-guided Prompt Learning for Vision-Language Models","date":"2023-06-01","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":3,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/prompt-pre-training-with-twenty-thousand-1","title":"Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition","date":"2023-04-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":5,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/towards-sustainable-self-supervised-learning","title":"Towards Sustainable Self-supervised Learning","date":"2022-10-20","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":6,"samples_unverified":2,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/maple-multi-modal-prompt-learning","title":"MaPLe: Multi-modal Prompt Learning","date":"2022-10-06","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pali-a-jointly-scaled-multilingual-language","title":"PaLI: A Jointly-Scaled Multilingual Language-Image Model","date":"2022-09-14","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/rf-next-efficient-receptive-field-search-for","title":"RF-Next: Efficient Receptive Field Search for Convolutional Neural Networks","date":"2022-06-14","rows_on_this_dataset":3,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/exploring-feature-self-relation-for-self","title":"SERE: Exploring Feature Self-relation for Self-supervised Transformer","date":"2022-06-10","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":1,"samples_unverified":2,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/conditional-prompt-learning-for-vision","title":"Conditional Prompt Learning for Vision-Language Models","date":"2022-03-10","rows_on_this_dataset":1,"code_links":12,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/a-convnet-for-the-2020s","title":"A ConvNet for the 2020s","date":"2022-01-10","rows_on_this_dataset":1,"code_links":54,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":80,"samples_ran":54,"samples_unverified":26,"pointer_only_for_licence":11,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/masked-autoencoders-are-scalable-vision","title":"Masked Autoencoders Are Scalable Vision Learners","date":"2021-11-11","rows_on_this_dataset":4,"code_links":58,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":137,"samples_ran":71,"samples_unverified":66,"pointer_only_for_licence":73,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/large-scale-unsupervised-semantic","title":"Large-scale Unsupervised Semantic Segmentation","date":"2021-06-06","rows_on_this_dataset":6,"code_links":3,"syntology":null},{"paper":"/paper/picie-unsupervised-semantic-segmentation","title":"PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering","date":"2021-03-30","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-transferable-visual-models-from","title":"Learning Transferable Visual Models From Natural Language Supervision","date":"2021-02-26","rows_on_this_dataset":1,"code_links":82,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":20,"samples_ran":16,"samples_unverified":4,"pointer_only_for_licence":16,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unsupervised-semantic-segmentation-by","title":"Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals","date":"2021-02-11","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/deep-clustering-for-unsupervised-learning-of","title":"Deep Clustering for Unsupervised Learning of Visual Features","date":"2018-07-15","rows_on_this_dataset":1,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":5,"samples_unverified":2,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":17,"samples_harvested":323,"samples_ran":187,"samples_unverified":136,"pointer_only_for_licence":121,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}