{"url":"/sota/fine-grained-image-classification-on-nabirds","task":{"name":"Fine-Grained Image Classification","url":"/task/fine-grained-image-classification","note":null},"dataset":{"name":"NABirds","url":"/dataset/nabirds"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"**Fine-Grained Image Classification** is a task in computer vision where the goal is to classify images into subcategories within a larger category. For example, classifying different species of birds or different types of flowers. This task is considered to be fine-grained because it requires the model to distinguish between subtle differences in visual appearance and patterns, making it more challenging than regular image classification tasks.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Looking for the Devil in the Details](https://arxiv.org/pdf/1903.06150v2.pdf) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Accuracy"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Accuracy":"higher"}},"counts":{"rows":30,"rows_with_code":21,"rows_with_paper_page":30,"rows_dated":30,"rows_using_additional_data":9},"rows":[{"rank_in_archive_order":1,"model":"MetaFormer\n(MetaFormer-2,384)","metrics":{"Accuracy":"93.0%"},"uses_additional_data":true,"paper_date":"2022-03-05","paper":"/paper/metaformer-a-unified-meta-framework-for-fine","paper_url":"https://arxiv.org/abs/2203.02751v1","paper_title":"MetaFormer: A Unified Meta Framework for Fine-Grained Recognition","code":"https://github.com/dqshuai/metaformer","n_code_links":2,"syntology":{"n_ran":1,"n_unverified":13,"n_samples":14,"n_pointer_only_licence":0}},{"rank_in_archive_order":2,"model":"HERBS","metrics":{"Accuracy":"93.0%"},"uses_additional_data":false,"paper_date":"2023-03-11","paper":"/paper/fine-grained-visual-classification-with-high-1","paper_url":"https://arxiv.org/abs/2303.06442v2","paper_title":"Fine-grained Visual Classification with High-temperature Refinement and Background Suppression","code":"https://github.com/chou141253/FGVC-HERBS","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"PIM","metrics":{"Accuracy":"92.8%"},"uses_additional_data":true,"paper_date":"2022-02-08","paper":"/paper/a-novel-plug-in-module-for-fine-grained-1","paper_url":"https://arxiv.org/abs/2202.03822v1","paper_title":"A Novel Plug-in Module for Fine-Grained Visual Classification","code":"https://github.com/chou141253/fgvc-pim","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"MPSA","metrics":{"Accuracy":"92.5%"},"uses_additional_data":false,"paper_date":"2024-08-16","paper":"/paper/multi-granularity-part-sampling-attention-for","paper_url":"https://ieeexplore.ieee.org/document/10638479","paper_title":"Multi-Granularity Part Sampling Attention for Fine-Grained Visual Classification","code":"https://github.com/mobulan/MPSA","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"ViT-NeT\n(SwinV2-B)","metrics":{"Accuracy":"92.5%"},"uses_additional_data":false,"paper_date":"2022-07-17","paper":"/paper/vit-net-interpretable-vision-transformers","paper_url":"https://proceedings.mlr.press/v162/kim22g/kim22g.pdf","paper_title":"ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder","code":"https://github.com/jumpsnack/ViT-NeT","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"CSQA-Net","metrics":{"Accuracy":"92.3%"},"uses_additional_data":false,"paper_date":"2024-03-15","paper":"/paper/context-semantic-quality-awareness-network","paper_url":"https://arxiv.org/abs/2403.10298v1","paper_title":"Context-Semantic Quality Awareness Network for Fine-Grained Visual Categorization","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":7,"model":"I2-HOFI","metrics":{"Accuracy":"92.12%"},"uses_additional_data":false,"paper_date":"2024-10-20","paper":"/paper/interweaving-insights-high-order-feature","paper_url":"https://link.springer.com/article/10.1007/s11263-024-02260-y","paper_title":"Interweaving Insights: High-Order Feature Interaction for Fine-Grained Visual Recognition","code":"https://github.com/arindam-1991/i2-hofi","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"MDCM","metrics":{"Accuracy":"92.0%"},"uses_additional_data":false,"paper_date":"2025-04-12","paper":"/paper/multi-scale-activation-refinement-and-1","paper_url":"https://arxiv.org/abs/2504.09215v1","paper_title":"Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":9,"model":"CGL","metrics":{"Accuracy":"91.7%"},"uses_additional_data":false,"paper_date":"2025-01-06","paper":"/paper/universal-fine-grained-visual-categorization","paper_url":"https://ieeexplore.ieee.org/document/10829548","paper_title":"Universal Fine-grained Visual Categorization by Concept Guided Learning","code":"https://github.com/biqiwhu/cgl","n_code_links":1,"syntology":null},{"rank_in_archive_order":10,"model":"SR-GNN","metrics":{"Accuracy":"91.2%"},"uses_additional_data":false,"paper_date":"2022-09-05","paper":"/paper/sr-gnn-spatial-relation-aware-graph-neural","paper_url":"https://arxiv.org/abs/2209.02109v1","paper_title":"SR-GNN: Spatial Relation-aware Graph Neural Network for Fine-Grained Image Categorization","code":"https://github.com/ardhendubehera/sr-gnn","n_code_links":1,"syntology":null},{"rank_in_archive_order":11,"model":"FAL-ViT","metrics":{"Accuracy":"91.1%"},"uses_additional_data":false,"paper_date":"2025-01-28","paper":"/paper/an-attention-locating-algorithm-for","paper_url":"https://ieeexplore.ieee.org/document/10855837","paper_title":"An Attention-Locating Algorithm for Eliminating Background Effects in Fine-grained Visual Classification","code":"https://github.com/yueting-huang/fal-vit","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"SFETrans","metrics":{"Accuracy":"91.1"},"uses_additional_data":false,"paper_date":"2025-06-14","paper":"/paper/structural-feature-enhanced-transformer-for","paper_url":"https://www.sciencedirect.com/science/article/abs/pii/S0031320325006156?dgcid=rss_sd_all","paper_title":"Structural feature enhanced transformer for fine-grained image recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":13,"model":"CAP","metrics":{"Accuracy":"91.0%"},"uses_additional_data":false,"paper_date":"2021-01-17","paper":"/paper/context-aware-attentional-pooling-cap-for","paper_url":"https://arxiv.org/abs/2101.06635v1","paper_title":"Context-aware Attentional Pooling (CAP) for Fine-grained Visual Classification","code":"https://github.com/ArdhenduBehera/cap","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":2,"n_samples":2,"n_pointer_only_licence":0}},{"rank_in_archive_order":14,"model":"MP-FGVC","metrics":{"Accuracy":"91.0%"},"uses_additional_data":false,"paper_date":"2023-09-16","paper":"/paper/delving-into-multimodal-prompting-for-fine","paper_url":"https://arxiv.org/abs/2309.08912v2","paper_title":"Delving into Multimodal Prompting for Fine-grained Visual Classification","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":15,"model":"TransIFC","metrics":{"Accuracy":"90.9%"},"uses_additional_data":false,"paper_date":"2022-12-31","paper":"/paper/transifc-invariant-cues-aware-feature","paper_url":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10023961","paper_title":"TransIFC: Invariant Cues-aware Feature Concentration Learning for Efficient Fine-grained Bird Image Classification","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":16,"model":"TransFG","metrics":{"Accuracy":"90.8%"},"uses_additional_data":false,"paper_date":"2021-03-14","paper":"/paper/transfg-a-transformer-architecture-for-fine","paper_url":"https://arxiv.org/abs/2103.07976v5","paper_title":"TransFG: A Transformer Architecture for Fine-grained Recognition","code":"https://github.com/TACJu/TransFG","n_code_links":2,"syntology":{"n_ran":3,"n_unverified":0,"n_samples":3,"n_pointer_only_licence":2}},{"rank_in_archive_order":17,"model":"IELT","metrics":{"Accuracy":"90.8%"},"uses_additional_data":false,"paper_date":"2023-02-13","paper":"/paper/fine-grained-visual-classification-via-2","paper_url":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10042971","paper_title":"Fine-Grained Visual Classification via Internal Ensemble Learning Transformer","code":"https://github.com/mobulan/ielt","n_code_links":1,"syntology":null},{"rank_in_archive_order":18,"model":"FVE","metrics":{"Accuracy":"90.3%"},"uses_additional_data":false,"paper_date":"2020-07-04","paper":"/paper/end-to-end-learning-of-a-fisher-vector","paper_url":"https://arxiv.org/abs/2007.02080v2","paper_title":"End-to-end Learning of a Fisher Vector Encoding for Part Features in Fine-grained Recognition","code":"https://github.com/DiKorsch/deep_fve","n_code_links":1,"syntology":null},{"rank_in_archive_order":19,"model":"TPSKG","metrics":{"Accuracy":"90.1%"},"uses_additional_data":false,"paper_date":"2021-07-14","paper":"/paper/transformer-with-peak-suppression-and","paper_url":"https://arxiv.org/abs/2107.06538v2","paper_title":"Transformer with Peak Suppression and Knowledge Guidance for Fine-grained Image Recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":20,"model":"FixSENet-154","metrics":{"Accuracy":"89.2%"},"uses_additional_data":true,"paper_date":"2019-06-14","paper":"/paper/fixing-the-train-test-resolution-discrepancy","paper_url":"https://arxiv.org/abs/1906.06423v4","paper_title":"Fixing the train-test resolution discrepancy","code":"https://github.com/facebookresearch/FixRes","n_code_links":3,"syntology":{"n_ran":0,"n_unverified":2,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":21,"model":"MGE-CNN","metrics":{"Accuracy":"88.6%"},"uses_additional_data":true,"paper_date":"2019-10-01","paper":"/paper/learning-a-mixture-of-granularity-specific","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Zhang_Learning_a_Mixture_of_Granularity-Specific_Experts_for_Fine-Grained_Categorization_ICCV_2019_paper.html","paper_title":"Learning a Mixture of Granularity-Specific Experts for Fine-Grained Categorization","code":"https://github.com/kalelpark/Mixture-of-Granularity-Specific-Experts-for-Fine-Grained-Categorization","n_code_links":1,"syntology":null},{"rank_in_archive_order":22,"model":"CS-Parts","metrics":{"Accuracy":"88.5%"},"uses_additional_data":true,"paper_date":"2019-09-16","paper":"/paper/classification-specific-parts-for-improving","paper_url":"https://arxiv.org/abs/1909.07075v1","paper_title":"Classification-Specific Parts for Improving Fine-Grained Visual Categorization","code":"https://github.com/DiKorsch/l1_parts","n_code_links":2,"syntology":null},{"rank_in_archive_order":23,"model":"CS-Part","metrics":{"Accuracy":"88.5%"},"uses_additional_data":false,"paper_date":"2019-09-16","paper":"/paper/classification-specific-parts-for-improving","paper_url":"https://arxiv.org/abs/1909.07075v1","paper_title":"Classification-Specific Parts for Improving Fine-Grained Visual Categorization","code":"https://github.com/DiKorsch/l1_parts","n_code_links":2,"syntology":null},{"rank_in_archive_order":24,"model":"API-Net","metrics":{"Accuracy":"88.1%"},"uses_additional_data":true,"paper_date":"2020-02-24","paper":"/paper/learning-attentive-pairwise-interaction-for","paper_url":"https://arxiv.org/abs/2002.10191v1","paper_title":"Learning Attentive Pairwise Interaction for Fine-Grained Classification","code":"https://github.com/PeiqinZhuang/API-Net","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":1,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":25,"model":"PAIRS","metrics":{"Accuracy":"87.9%"},"uses_additional_data":true,"paper_date":"2018-01-27","paper":"/paper/aligned-to-the-object-not-to-the-image-a","paper_url":"http://arxiv.org/abs/1801.09057v4","paper_title":"Aligned to the Object, not to the Image: A Unified Pose-aligned Representation for Fine-grained Recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":26,"model":"Cross-X","metrics":{"Accuracy":"86.4%"},"uses_additional_data":true,"paper_date":"2019-09-10","paper":"/paper/cross-x-learning-for-fine-grained-visual","paper_url":"https://arxiv.org/abs/1909.04412v1","paper_title":"Cross-X Learning for Fine-Grained Visual Categorization","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":27,"model":"MaxEnt-CNN","metrics":{"Accuracy":"83.0%"},"uses_additional_data":true,"paper_date":"2018-12-01","paper":"/paper/maximum-entropy-fine-grained-classification","paper_url":"http://papers.nips.cc/paper/7344-maximum-entropy-fine-grained-classification","paper_title":"Maximum-Entropy Fine Grained Classification","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":28,"model":"PC-DenseNet-161","metrics":{"Accuracy":"82.79%"},"uses_additional_data":false,"paper_date":"2017-05-22","paper":"/paper/pairwise-confusion-for-fine-grained-visual","paper_url":"http://arxiv.org/abs/1705.08016v3","paper_title":"Pairwise Confusion for Fine-Grained Visual Classification","code":"https://github.com/abhimanyudubey/confusion","n_code_links":1,"syntology":null},{"rank_in_archive_order":29,"model":"BYOL+CVSA (ResNet-50)","metrics":{"Accuracy":"79.64%"},"uses_additional_data":false,"paper_date":"2021-06-30","paper":"/paper/align-yourself-self-supervised-pre-training","paper_url":"https://arxiv.org/abs/2106.15788v4","paper_title":"Exploring Localization for Self-supervised Fine-grained Contrastive Learning","code":"https://github.com/Westlake-AI/openmixup","n_code_links":1,"syntology":null},{"rank_in_archive_order":30,"model":"Bilinear-CNN","metrics":{"Accuracy":"79.4%"},"uses_additional_data":false,"paper_date":"2015-04-29","paper":"/paper/bilinear-cnns-for-fine-grained-visual","paper_url":"http://arxiv.org/abs/1504.07889v6","paper_title":"Bilinear CNNs for Fine-grained Visual Recognition","code":"https://github.com/tommarvoloriddle/Bilinear-CNN-Tensorflow2.4-implementation","n_code_links":4,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9581,"papers_checked":6264,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":3316},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":5,"rows_with_any_sample_ran":3,"distinct_papers_with_graph_line":5,"distinct_papers_with_any_sample_ran":3,"samples_over_distinct_papers":{"n_ran":5,"n_unverified":18,"n_samples":23,"n_pointer_only_licence":6,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":5,"n_unverified":18,"n_samples":23,"n_pointer_only_licence":6,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}