{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/aligned-to-the-object-not-to-the-image-a","title":"Aligned to the Object, not to the Image: A Unified Pose-aligned Representation for Fine-grained Recognition","arxiv_id":"1801.09057","date":"2018-01-27","proceeding":null,"authors":["Pei Guo","Ryan Farrell"],"abstract":"Dramatic appearance variation due to pose constitutes a great challenge in\nfine-grained recognition, one which recent methods using attention mechanisms\nor second-order statistics fail to adequately address. Modern CNNs typically\nlack an explicit understanding of object pose and are instead confused by\nentangled pose and appearance. In this paper, we propose a unified object\nrepresentation built from a hierarchy of pose-aligned regions. Rather than\nrepresenting an object by regions aligned to image axes, the proposed\nrepresentation characterizes appearance relative to the object's pose using\npose-aligned patches whose features are robust to variations in pose, scale and\nrotation. We propose an algorithm that performs pose estimation and forms the\nunified object representation as the concatenation of hierarchical pose-aligned\nregions features, which is then fed into a classification network. The proposed\nalgorithm surpasses the performance of other approaches, increasing the\nstate-of-the-art by nearly 2% on the widely-used CUB-200 dataset and by more\nthan 8% on the much larger NABirds dataset. The effectiveness of this paradigm\nrelative to competing methods suggests the critical importance of disentangling\npose and appearance for continued progress in fine-grained recognition.","url_abs":"http://arxiv.org/abs/1801.09057v4","url_pdf":"http://arxiv.org/pdf/1801.09057v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"object","task_name":"Object"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"PAIRS","rank_in_archive_order":25,"of":30,"metrics":{"Accuracy":"87.9%"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.09057","atlas_url":"https://app.syntology.ai/?focus=1801.09057","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}