{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-object-frames-by","title":"Unsupervised learning of object frames by dense equivariant image labelling","arxiv_id":"1706.02932","date":"2017-06-09","proceeding":"NeurIPS 2017 12","authors":["James Thewlis","Hakan Bilen","Andrea Vedaldi"],"abstract":"One of the key challenges of visual perception is to extract abstract models\nof 3D objects and object categories from visual measurements, which are\naffected by complex nuisance factors such as viewpoint, occlusion, motion, and\ndeformations. Starting from the recent idea of viewpoint factorization, we\npropose a new approach that, given a large number of images of an object and no\nother supervision, can extract a dense object-centric coordinate frame. This\ncoordinate frame is invariant to deformations of the images and comes with a\ndense equivariant labelling neural network that can map image pixels to their\ncorresponding object coordinates. We demonstrate the applicability of this\nmethod to simple articulated objects and deformable objects such as human\nfaces, learning embeddings from random synthetic transformations or optical\nflow correspondences, all without any manual supervision.","url_abs":"http://arxiv.org/abs/1706.02932v2","url_pdf":"http://arxiv.org/pdf/1706.02932v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"unsupervised-facial-landmark-detection","task_name":"Unsupervised Facial Landmark Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on","task":"Unsupervised Facial Landmark Detection","dataset":"300W","model":"DEIL","rank_in_archive_order":4,"of":4,"metrics":{"NME":"8.23"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-2","task":"Unsupervised Facial Landmark Detection","dataset":"AFLW (Zhang CVPR 2018 crops)","model":"DEIL","rank_in_archive_order":4,"of":4,"metrics":{"NME":"10.14"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-3","task":"Unsupervised Facial Landmark Detection","dataset":"AFLW-MTFL","model":"DEIL","rank_in_archive_order":3,"of":3,"metrics":{"NME":"10.99"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-facial-landmark-detection-on-1","task":"Unsupervised Facial Landmark Detection","dataset":"MAFL","model":"DEIL","rank_in_archive_order":8,"of":13,"metrics":{"NME":"4.02"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.02932","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}