{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-object-semantic","title":"Unsupervised learning of object semantic parts from internal states of CNNs by population encoding","arxiv_id":"1511.06855","date":"2015-11-21","proceeding":null,"authors":["Jianyu Wang","Zhishuai Zhang","Cihang Xie","Vittal Premachandran","Alan Yuille"],"abstract":"We address the key question of how object part representations can be found\nfrom the internal states of CNNs that are trained for high-level tasks, such as\nobject classification. This work provides a new unsupervised method to learn\nsemantic parts and gives new understanding of the internal representations of\nCNNs. Our technique is based on the hypothesis that semantic parts are\nrepresented by populations of neurons rather than by single filters. We propose\na clustering technique to extract part representations, which we call Visual\nConcepts. We show that visual concepts are semantically coherent in that they\nrepresent semantic parts, and visually coherent in that corresponding image\npatches appear very similar. Also, visual concepts provide full spatial\ncoverage of the parts of an object, rather than a few sparse parts as is\ntypically found in keypoint annotations. Furthermore, We treat single visual\nconcept as part detector and evaluate it for keypoint detection using the\nPASCAL3D+ dataset and for part detection using our newly annotated ImageNetPart\ndataset. The experiments demonstrate that visual concepts can be used to detect\nparts. We also show that some visual concepts respond to several semantic\nparts, provided these parts are visually similar. Thus visual concepts have the\nessential properties: semantic meaning and detection capability. Note that our\nImageNetPart dataset gives rich part annotations which cover the whole object,\nmaking it useful for other part-related applications.","url_abs":"http://arxiv.org/abs/1511.06855v3","url_pdf":"http://arxiv.org/pdf/1511.06855v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-object-semantic","repo_url":"https://github.com/ytongbai/SemanticPartDetection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"keypoint-detection","task_name":"Keypoint Detection"},{"task_slug":"object","task_name":"Object"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1511.06855","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}