{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/large-scale-unsupervised-semantic","title":"Large-scale Unsupervised Semantic Segmentation","arxiv_id":"2106.03149","date":"2021-06-06","proceeding":null,"authors":["ShangHua Gao","Zhong-Yu Li","Ming-Hsuan Yang","Ming-Ming Cheng","Junwei Han","Philip Torr"],"abstract":"Empowered by large datasets, e.g., ImageNet, unsupervised learning on large-scale data has enabled significant advances for classification tasks. However, whether the large-scale unsupervised semantic segmentation can be achieved remains unknown. There are two major challenges: i) we need a large-scale benchmark for assessing algorithms; ii) we need to develop methods to simultaneously learn category and shape representation in an unsupervised manner. In this work, we propose a new problem of large-scale unsupervised semantic segmentation (LUSS) with a newly created benchmark dataset to help the research progress. Building on the ImageNet dataset, we propose the ImageNet-S dataset with 1.2 million training images and 50k high-quality semantic segmentation annotations for evaluation. Our benchmark has a high data diversity and a clear task objective. We also present a simple yet effective method that works surprisingly well for LUSS. In addition, we benchmark related un/weakly/fully supervised methods accordingly, identifying the challenges and possible directions of LUSS. The benchmark and source code is publicly available at https://github.com/LUSSeg.","url_abs":"https://arxiv.org/abs/2106.03149v3","url_pdf":"https://arxiv.org/pdf/2106.03149v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"large-scale-unsupervised-semantic","repo_url":"https://github.com/LUSSeg/ImageNet-S","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"large-scale-unsupervised-semantic","repo_url":"https://github.com/LUSSeg/ImageNetSegModel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"large-scale-unsupervised-semantic","repo_url":"https://github.com/LUSSeg/PASS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-semantic-segmentation","task_name":"Unsupervised Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"imagenet-s","name":"ImageNet-S","full_name":"ImageNet Semantic Segmentation"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-imagenet-s","task":"Semantic Segmentation","dataset":"ImageNet-S","model":"PASS (ResNet-50 D16, 224x224, LUSS)","rank_in_archive_order":19,"of":20,"metrics":{"mIoU (test)":"20.8","mIoU (val)":"21.6"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-imagenet-s","task":"Semantic Segmentation","dataset":"ImageNet-S","model":"PASS (ResNet-50 D32, 224x224, LUSS)","rank_in_archive_order":20,"of":20,"metrics":{"mIoU (test)":"20.3","mIoU (val)":"21.0"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-4","task":"Unsupervised Semantic Segmentation","dataset":"ImageNet-S","model":"PASS","rank_in_archive_order":1,"of":1,"metrics":{"mIoU (test)":"11.0","mIoU (val)":"11.5"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-5","task":"Unsupervised Semantic Segmentation","dataset":"ImageNet-S-300","model":"PASS","rank_in_archive_order":1,"of":1,"metrics":{"mIoU (test)":"18.1","mIoU (val)":"18"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-6","task":"Unsupervised Semantic Segmentation","dataset":"ImageNet-S-50","model":"PASS (+Saliency map)","rank_in_archive_order":1,"of":5,"metrics":{"mIoU (test)":"42.3","mIoU (val)":"43.3"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-semantic-segmentation-on-6","task":"Unsupervised Semantic Segmentation","dataset":"ImageNet-S-50","model":"PASS","rank_in_archive_order":2,"of":5,"metrics":{"mIoU (test)":"32.0","mIoU (val)":"32.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2106.03149","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}