{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transifc-invariant-cues-aware-feature","title":"TransIFC: Invariant Cues-aware Feature Concentration Learning for Efficient Fine-grained Bird Image Classification","arxiv_id":null,"date":"2022-12-31","proceeding":"TIP 2022 12","authors":["Hai Liu","Cheng Zhang","Yongjian Deng","Bochen Xie","Tingting Liu","Zhaoli Zhang","You-Fu Li"],"abstract":"Fine-grained bird image classification (FBIC) is not\r\nonly meaningful for endangered bird observation and protection\r\nbut also a prevalent task for image classification in multimedia\r\nprocessing and computer vision. However, FBIC suffers from\r\nseveral challenges, such as bird molting, complex background, and\r\narbitrary bird posture. To effectively tackle these challenges, we\r\npresent a novel invariant cues-aware feature concentration\r\nTransformer (TransIFC), which learns invariant and core\r\ninformation in bird images. To this end, two novel modules are\r\nproposed to leverage the characteristics of bird images, namely,\r\nthe hierarchy stage feature aggregation (HSFA) module and the\r\nfeature in feature abstraction (FFA) module. The HSFA module\r\naggregates the multiscale information of bird images by\r\nconcatenating multilayer features. The FFA module extracts the\r\ninvariant cues of birds through feature selection based on\r\ndiscrimination scores. Transformer is employed as the backbone\r\nto reveal the long-dependent semantic relationships in bird\r\nimages. Moreover, abundant visualizations are provided to prove\r\nthe interpretability of the HSFA and FFA modules in TransIFC.\r\nComprehensive experiments demonstrate that TransIFC can\r\nachieve state-of-the-art performance on the CUB-200-2011 dataset\r\n(91.0%) and the NABirds dataset (90.9%). Finally, extended\r\nexperiments have been conducted on the Stanford Cars dataset to\r\nsuggest the potential of generalizing our method on other finegrained visual classification tasks.","url_abs":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10023961","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10023961","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"feature-selection","task_name":"feature selection"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"feature-selection","method_name":"Feature Selection"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"TransIFC","rank_in_archive_order":15,"of":30,"metrics":{"Accuracy":"90.9%"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}