{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-single-noun-prior-for-image-clustering","title":"Dataset Summarization by K Principal Concepts","arxiv_id":"2104.03952","date":"2021-04-08","proceeding":null,"authors":["Niv Cohen","Yedid Hoshen"],"abstract":"We propose the new task of K principal concept identification for dataset summarizarion. The objective is to find a set of K concepts that best explain the variation within the dataset. Concepts are high-level human interpretable terms such as \"tiger\", \"kayaking\" or \"happy\". The K concepts are selected from a (potentially long) input list of candidates, which we denote the concept-bank. The concept-bank may be taken from a generic dictionary or constructed by task-specific prior knowledge. An image-language embedding method (e.g. CLIP) is used to map the images and the concept-bank into a shared feature space. To select the K concepts that best explain the data, we formulate our problem as a K-uncapacitated facility location problem. An efficient optimization technique is used to scale the local search algorithm to very large concept-banks. The output of our method is a set of K principal concepts that summarize the dataset. Our approach provides a more explicit summary in comparison to selecting K representative images, which are often ambiguous. As a further application of our method, the K principal concepts can be used to classify the dataset into K groups. Extensive experiments demonstrate the efficacy of our approach.","url_abs":"https://arxiv.org/abs/2104.03952v2","url_pdf":"https://arxiv.org/pdf/2104.03952v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"image-clustering","task_name":"Image Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-clustering-on-cifar-10","task":"Image Clustering","dataset":"CIFAR-10","model":"Single-Noun Prior","rank_in_archive_order":21,"of":40,"metrics":{"ARI":"0.702","Accuracy":"0.853","Backbone":"ViT-B-32","NMI":"0.731","Train set":"Train+Test"},"uses_additional_data":true},{"leaderboard":"/sota/image-clustering-on-imagenet-100","task":"Image Clustering","dataset":"ImageNet-100 (TEMI Split)","model":"Single-Noun Prior","rank_in_archive_order":4,"of":5,"metrics":{"ACCURACY":"0.731","ARI":"0.628","NMI":"0.805"},"uses_additional_data":true},{"leaderboard":"/sota/image-clustering-on-imagenet-200","task":"Image Clustering","dataset":"ImageNet-200","model":"Single-Noun Prior","rank_in_archive_order":5,"of":5,"metrics":{"\t ACCURACY":"0.598","ARI":"0.486","NMI":"0.749"},"uses_additional_data":true},{"leaderboard":"/sota/image-clustering-on-imagenet-50-1","task":"Image Clustering","dataset":"ImageNet-50 (TEMI Split)","model":"Single-Noun Prior","rank_in_archive_order":4,"of":5,"metrics":{"ACCURACY":"0.827","ARI":"0.744","NMI":"0.847"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.03952","atlas_url":"https://app.syntology.ai/?focus=2104.03952","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}