{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/structural-feature-enhanced-transformer-for","title":"Structural feature enhanced transformer for fine-grained image recognition","arxiv_id":null,"date":"2025-06-14","proceeding":"Pattern Recognition 2025 6","authors":["Ying Yu","Wei Wei","Cairong Zhao","Jin Qian","Enhong Chen"],"abstract":"Existing fine-grained image recognition (FGIR) models mainly rely on high-level semantic features to extract discriminative information, ignoring the potential role of the overall structural information of objects and the structural relationships between key parts. To address this issue, we propose the Structural Feature Enhancement Transformer (SFETrans). SFETrans consists of a visual transformer backbone network responsible for extracting complex semantic features. Additionally, it includes a structural modeling (SM) branch and an amplitude component exchange (ACE) module, both dedicated to enhancing the learning of structural features. The SM branch actively models the structural relationships between key parts of objects and extracts corresponding structural features, while the ACE module guides the model to learn structural information in the phase spectrum by introducing implicit constraints during training. By synergizing the backbone network and the two modules, SFETrans exhibits competitive performance on four benchmark datasets and outperforms other comparison methods in terms of computational efficiency.","url_abs":"https://www.sciencedirect.com/science/article/abs/pii/S0031320325006156?dgcid=rss_sd_all","url_pdf":"https://www.sciencedirect.com/science/article/abs/pii/S0031320325006156?dgcid=rss_sd_all","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"fine-grained-image-recognition","task_name":"Fine-Grained Image Recognition"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-cub-200-1","task":"Fine-Grained Image Classification","dataset":"CUB-200-2011","model":"SFETrans","rank_in_archive_order":6,"of":30,"metrics":{"Accuracy":"91.8"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"SFETrans","rank_in_archive_order":12,"of":30,"metrics":{"Accuracy":"91.1"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-stanford-1","task":"Fine-Grained Image Classification","dataset":"Stanford Dogs","model":"SFETrans","rank_in_archive_order":10,"of":24,"metrics":{"Accuracy":"92.4"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}