{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-discovery-of-object-landmarks-as","title":"Unsupervised Discovery of Object Landmarks as Structural Representations","arxiv_id":"1804.04412","date":"2018-04-12","proceeding":"CVPR 2018 6","authors":["Yuting Zhang","Yijie Guo","Yixin Jin","Yijun Luo","Zhiyuan He","Honglak Lee"],"abstract":"Deep neural networks can model images with rich latent representations, but\nthey cannot naturally conceptualize structures of object categories in a\nhuman-perceptible way. This paper addresses the problem of learning object\nstructures in an image modeling process without supervision. We propose an\nautoencoding formulation to discover landmarks as explicit structural\nrepresentations. The encoding module outputs landmark coordinates, whose\nvalidity is ensured by constraints that reflect the necessary properties for\nlandmarks. The decoding module takes the landmarks as a part of the learnable\ninput representations in an end-to-end differentiable framework. Our discovered\nlandmarks are semantically meaningful and more predictive of manually annotated\nlandmarks than those discovered by previous methods. The coordinates of our\nlandmarks are also complementary features to pretrained deep-neural-network\nrepresentations in recognizing visual attributes. In addition, the proposed\nmethod naturally creates an unsupervised, perceptible interface to manipulate\nobject shapes and decode images with controllable structures. The project\nwebpage is at http://ytzhang.net/projects/lmdis-rep","url_abs":"http://arxiv.org/abs/1804.04412v1","url_pdf":"http://arxiv.org/pdf/1804.04412v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-discovery-of-object-landmarks-as","repo_url":"https://github.com/YutingZhang/lmdis-rep","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"unsupervised-facial-landmark-detection","task_name":"Unsupervised Facial Landmark Detection"},{"task_slug":"unsupervised-human-pose-estimation","task_name":"Unsupervised Human Pose Estimation"},{"task_slug":"unsupervised-keypoint-estimation","task_name":"Unsupervised Keypoint Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-2","task":"Unsupervised Facial Landmark Detection","dataset":"AFLW (Zhang CVPR 2018 crops)","model":"LMDIS-REP","rank_in_archive_order":3,"of":4,"metrics":{"NME":"6.58"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-1","task":"Unsupervised Facial Landmark Detection","dataset":"MAFL","model":"LMDIS-REP","rank_in_archive_order":4,"of":13,"metrics":{"NME":"3.15"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-5","task":"Unsupervised Facial Landmark Detection","dataset":"MAFL Unaligned","model":"LMDIS-REP","rank_in_archive_order":9,"of":9,"metrics":{"NME":"40.82"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.04412","atlas_url":"https://app.syntology.ai/?focus=1804.04412","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1804.04412"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/YutingZhang/lmdis-rep","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9ba5d8704cc281c4","entry":"get2d_deconv_output_size","repo":"YutingZhang/lmdis-rep","repo_kind":"listed","path":"net_modules/deconv.py","file_url":"https://github.com/YutingZhang/lmdis-rep/blob/HEAD/net_modules/deconv.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9ba5d8704cc281c4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}