{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-x-learning-for-fine-grained-visual","title":"Cross-X Learning for Fine-Grained Visual Categorization","arxiv_id":"1909.04412","date":"2019-09-10","proceeding":"ICCV 2019 10","authors":["Wei Luo","Xitong Yang","Xianjie Mo","Yuheng Lu","Larry S. Davis","Jun Li","Jian Yang","Ser-Nam Lim"],"abstract":"Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features are extracted for fine-grained classification. However, these methods typically treat the part-specific features of each image in isolation while neglecting their relationships between different images. In this paper, we propose Cross-X learning, a simple yet effective approach that exploits the relationships between different images and between different network layers for robust multi-scale feature learning. Our approach involves two novel components: (i) a cross-category cross-semantic regularizer that guides the extracted features to represent semantic parts and, (ii) a cross-layer regularizer that improves the robustness of multi-scale features by matching the prediction distribution across multiple layers. Our approach can be easily trained end-to-end and is scalable to large datasets like NABirds. We empirically analyze the contributions of different components of our approach and demonstrate its robustness, effectiveness and state-of-the-art performance on five benchmark datasets. Code is available at \\url{https://github.com/cswluo/CrossX}.","url_abs":"https://arxiv.org/abs/1909.04412v1","url_pdf":"https://arxiv.org/pdf/1909.04412v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"fine-grained-visual-categorization","task_name":"Fine-Grained Visual Categorization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-cub-200","task":"Fine-Grained Image Classification","dataset":"CUB-200-2011","model":"Cross-X","rank_in_archive_order":10,"of":12,"metrics":{"Accuracy":"87.7%"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-fgvc","task":"Fine-Grained Image Classification","dataset":"FGVC Aircraft","model":"Cross-X","rank_in_archive_order":37,"of":57,"metrics":{"Accuracy":"92.7%"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"Cross-X","rank_in_archive_order":26,"of":30,"metrics":{"Accuracy":"86.4%"},"uses_additional_data":true},{"leaderboard":"/sota/fine-grained-image-classification-on-stanford","task":"Fine-Grained Image Classification","dataset":"Stanford Cars","model":"Cross-X","rank_in_archive_order":38,"of":83,"metrics":{"Accuracy":"94.6%"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1909.04412","atlas_url":"https://app.syntology.ai/?focus=1909.04412","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}