{"url":"/sota/fine-grained-image-classification-on-caltech","task":{"name":"Fine-Grained Image Classification","url":"/task/fine-grained-image-classification","note":null},"dataset":{"name":"Caltech-101","url":"/dataset/caltech-101"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"**Fine-Grained Image Classification** is a task in computer vision where the goal is to classify images into subcategories within a larger category. For example, classifying different species of birds or different types of flowers. This task is considered to be fine-grained because it requires the model to distinguish between subtle differences in visual appearance and patterns, making it more challenging than regular image classification tasks.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Looking for the Devil in the Details](https://arxiv.org/pdf/1903.06150v2.pdf) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Top-1 Error Rate","Accuracy"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Top-1 Error Rate":null,"Accuracy":"higher"}},"counts":{"rows":18,"rows_with_code":15,"rows_with_paper_page":18,"rows_dated":18,"rows_using_additional_data":5},"rows":[{"rank_in_archive_order":1,"model":"VIT-L/16","metrics":{"Top-1 Error Rate":"1.98%"},"uses_additional_data":false,"paper_date":"2023-05-05","paper":"/paper/reduction-of-class-activation-uncertainty","paper_url":"https://arxiv.org/abs/2305.03238v6","paper_title":"Reduction of Class Activation Uncertainty with Background Information","code":"https://github.com/dipuk0506/SpinalNet","n_code_links":2,"syntology":null},{"rank_in_archive_order":2,"model":"Wide-ResNet-101 (Spinal FC)","metrics":{"Accuracy":"97.32","Top-1 Error Rate":"2.68%"},"uses_additional_data":true,"paper_date":"2020-07-07","paper":"/paper/spinalnet-deep-neural-network-with-gradual-1","paper_url":"https://arxiv.org/abs/2007.03347v3","paper_title":"SpinalNet: Deep Neural Network with Gradual Input","code":"https://github.com/dipuk0506/SpinalNet","n_code_links":3,"syntology":null},{"rank_in_archive_order":3,"model":"Wide-ResNet-101","metrics":{"Top-1 Error Rate":"2.89%"},"uses_additional_data":true,"paper_date":"2020-07-07","paper":"/paper/spinalnet-deep-neural-network-with-gradual-1","paper_url":"https://arxiv.org/abs/2007.03347v3","paper_title":"SpinalNet: Deep Neural Network with Gradual Input","code":"https://github.com/dipuk0506/SpinalNet","n_code_links":3,"syntology":null},{"rank_in_archive_order":4,"model":"µ2Net+ (ViT-L/16)","metrics":{"Top-1 Error Rate":"4.06%"},"uses_additional_data":false,"paper_date":"2022-09-15","paper":"/paper/a-continual-development-methodology-for-large","paper_url":"https://arxiv.org/abs/2209.07326v3","paper_title":"A Continual Development Methodology for Large-scale Multitask Dynamic ML Systems","code":"https://github.com/google-research/google-research/tree/master/muNet","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"ResNeXt-101-32x8d","metrics":{"Accuracy":"95.58","Top-1 Error Rate":"4.42%"},"uses_additional_data":false,"paper_date":"2021-08-31","paper":"/paper/dead-pixel-test-using-effective-receptive","paper_url":"https://arxiv.org/abs/2108.13576v1","paper_title":"Dead Pixel Test Using Effective Receptive Field","code":"https://github.com/kmbmjn/deadpixeltest_code","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"TWIST (ResNet-50 )","metrics":{"Accuracy":"93.5%","Top-1 Error Rate":"6.5%"},"uses_additional_data":false,"paper_date":"2021-10-14","paper":"/paper/self-supervised-learning-by-estimating-twin-1","paper_url":"https://arxiv.org/abs/2110.07402v4","paper_title":"Self-Supervised Learning by Estimating Twin Class Distributions","code":"https://github.com/bytedance/TWIST","n_code_links":2,"syntology":{"n_ran":5,"n_unverified":10,"n_samples":15,"n_pointer_only_licence":0}},{"rank_in_archive_order":7,"model":"VGG-19bn (Spinal FC)","metrics":{"Top-1 Error Rate":"6.84%"},"uses_additional_data":true,"paper_date":"2020-07-07","paper":"/paper/spinalnet-deep-neural-network-with-gradual-1","paper_url":"https://arxiv.org/abs/2007.03347v3","paper_title":"SpinalNet: Deep Neural Network with Gradual Input","code":"https://github.com/dipuk0506/SpinalNet","n_code_links":3,"syntology":null},{"rank_in_archive_order":8,"model":"µ2Net (ViT-L/16)","metrics":{"Top-1 Error Rate":"7%"},"uses_additional_data":false,"paper_date":"2022-05-25","paper":"/paper/an-evolutionary-approach-to-dynamic","paper_url":"https://arxiv.org/abs/2205.12755v6","paper_title":"An Evolutionary Approach to Dynamic Introduction of Tasks in Large-scale Multitask Learning Systems","code":"https://github.com/google-research/google-research/tree/master/muNet","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"NNCLR","metrics":{"Top-1 Error Rate":"8.7%"},"uses_additional_data":false,"paper_date":"2021-04-29","paper":"/paper/with-a-little-help-from-my-friends-nearest","paper_url":"https://arxiv.org/abs/2104.14548v2","paper_title":"With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations","code":"https://github.com/lightly-ai/lightly","n_code_links":4,"syntology":{"n_ran":4,"n_unverified":1,"n_samples":5,"n_pointer_only_licence":0}},{"rank_in_archive_order":10,"model":"SEER (RegNet10B - linear eval)","metrics":{"Accuracy":"91.0","Top-1 Error Rate":"9.0%"},"uses_additional_data":true,"paper_date":"2022-02-16","paper":"/paper/vision-models-are-more-robust-and-fair-when","paper_url":"https://arxiv.org/abs/2202.08360v2","paper_title":"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision","code":"https://github.com/facebookresearch/vissl","n_code_links":1,"syntology":null},{"rank_in_archive_order":11,"model":"ViT-S/16 (RPE w/ GAB)","metrics":{"Top-1 Error Rate":"9.798%"},"uses_additional_data":false,"paper_date":"2023-05-08","paper":"/paper/understanding-gaussian-attention-bias-of","paper_url":"https://arxiv.org/abs/2305.04722v1","paper_title":"Understanding Gaussian Attention Bias of Vision Transformers Using Effective Receptive Fields","code":"https://github.com/kmbmjn/GaussianAttentionBias","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"AutoAugment","metrics":{"Top-1 Error Rate":"13.07%"},"uses_additional_data":false,"paper_date":"2018-05-24","paper":"/paper/autoaugment-learning-augmentation-policies","paper_url":"http://arxiv.org/abs/1805.09501v3","paper_title":"AutoAugment: Learning Augmentation Policies from Data","code":"https://github.com/tensorflow/models/tree/master/research/autoaugment","n_code_links":33,"syntology":{"n_ran":6,"n_unverified":37,"n_samples":43,"n_pointer_only_licence":2}},{"rank_in_archive_order":13,"model":"PreResNet-101","metrics":{"Top-1 Error Rate":"15.8036%"},"uses_additional_data":false,"paper_date":"2023-02-13","paper":"/paper/how-to-use-dropout-correctly-on-residual","paper_url":"https://arxiv.org/abs/2302.06112v1","paper_title":"How to Use Dropout Correctly on Residual Networks with Batch Normalization","code":"https://github.com/kmbmjn/DropoutCorrectly","n_code_links":1,"syntology":null},{"rank_in_archive_order":14,"model":"SE-ResNet-101 (SAP)","metrics":{"Top-1 Error Rate":"15.949%"},"uses_additional_data":false,"paper_date":"2024-09-25","paper":"/paper/stochastic-subsampling-with-average-pooling","paper_url":"https://arxiv.org/abs/2409.16630v1","paper_title":"Stochastic Subsampling With Average Pooling","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":15,"model":"ResNet-101 (ideal number of groups)","metrics":{"Top-1 Error Rate":"22.247%"},"uses_additional_data":false,"paper_date":"2023-02-07","paper":"/paper/on-the-ideal-number-of-groups-for-isometric","paper_url":"https://arxiv.org/abs/2302.03193v1","paper_title":"On the Ideal Number of Groups for Isometric Gradient Propagation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":16,"model":"UL-Hopfield (ULH)","metrics":{"Accuracy":"91.00"},"uses_additional_data":false,"paper_date":"2018-05-02","paper":"/paper/unsupervised-learning-using-pretrained-cnn","paper_url":"http://arxiv.org/abs/1805.01033v1","paper_title":"Unsupervised Learning using Pretrained CNN and Associative Memory Bank","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":17,"model":"Bamboo (ViT-B/16)","metrics":{"Accuracy":"94.8"},"uses_additional_data":true,"paper_date":"2022-03-15","paper":"/paper/bamboo-building-mega-scale-vision-dataset","paper_url":"https://arxiv.org/abs/2203.07845v2","paper_title":"Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy","code":"https://github.com/zhangyuanhan-ai/bamboo","n_code_links":2,"syntology":{"n_ran":3,"n_unverified":0,"n_samples":3,"n_pointer_only_licence":3}},{"rank_in_archive_order":18,"model":"Pre trained wide-resnet-101","metrics":{"Accuracy":"97.76"},"uses_additional_data":false,"paper_date":"2021-03-21","paper":"/paper/progressivespinalnet-architecture-for-fc","paper_url":"https://arxiv.org/abs/2103.11373v1","paper_title":"ProgressiveSpinalNet architecture for FC layers","code":"https://github.com/praveenchopra/ProgressiveSpinalNet","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":4,"rows_with_any_sample_ran":4,"distinct_papers_with_graph_line":4,"distinct_papers_with_any_sample_ran":4,"samples_over_distinct_papers":{"n_ran":18,"n_unverified":48,"n_samples":66,"n_pointer_only_licence":5,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":18,"n_unverified":48,"n_samples":66,"n_pointer_only_licence":5,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}