Browse State-of-the-Art › Fine-Grained Image Classification
Fine-Grained Image Classification
200 papers with code · 36 benchmarks · 41 datasets archive 2025-07-28
Fine-Grained Image Classification is a task in computer vision where the goal is to classify images into subcategories within a larger category. For example, classifying different species of birds or different types of flowers. This task is considered to be fine-grained because it requires the model to distinguish between subtle differences in visual appearance and patterns, making it more challenging than regular image classification tasks.
( Image credit: Looking for the Devil in the Details )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
36 leaderboard tables shown for this task, 36 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 36 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
41 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 41 until expanded.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 200 papers with code (353 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 May 2019 144 repositories listed Syntology ran 171 of 302 samples · 131 unverified · 112 pointer-only (licence)Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available.
-
23 Dec 2020 40 repositories listed Syntology ran 12 of 19 samples · 7 unverified · 3 pointer-only (licence)In this work, we produce a competitive convolution-free transformer by training on Imagenet only.
-
24 May 2018 33 repositories listed Syntology ran 6 of 43 samples · 37 unverified · 2 pointer-only (licence)In our implementation, we have designed a search space where a policy consists of many sub-policies, one of which is randomly chosen for each image in each mini-batch.
-
14 Apr 2023 26 repositories listed Syntology ran 21 of 46 samples · 25 unverified · 12 pointer-only (licence)The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision.
-
7 May 2021 19 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe present ResMLP, an architecture built entirely upon multi-layer perceptrons for image classification.
-
3 Oct 2020 18 repositories listed Syntology ran 8 of 20 samples · 12 unverified · 6 pointer-only (licence)In today's heavily overparameterized models, the value of the training loss provides few guarantees on model generalization ability.
-
1 Oct 2021 14 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedWe share competitive training settings and pre-trained models in the timm open-source library, with the hope that they will serve as better baselines for future work.
-
16 Nov 2018 13 repositories listed Syntology ran 1 of 25 samples · 24 unverified · 16 pointer-only (licence)Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks.
-
27 Feb 2021 12 repositories listed Syntology ran 16 of 24 samples · 8 unverified · 5 pointer-only (licence)In this paper, we point out that the attention inside these local patches are also essential for building visual transformers with high performance and we explore a new architecture, namely, Transformer iN Transformer…
-
2 Sep 2018 12 repositories listed Syntology ran 3 of 9 samples · 6 unverifiedIn consideration of intrinsic consistency between informativeness of the regions and their probability being ground-truth class, we design a novel training paradigm, which enables Navigator to detect most informative…
-
12 Apr 2021 9 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 1 pointer-only (licence)Our models are flexible in terms of model size, and can have as little as 0.
-
24 Dec 2019 9 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedWe conduct detailed analysis of the main components that lead to high transfer performance.
-
18 Mar 2022 8 repositories listed Syntology ran 1 of 1 samples · 0 unverified(2) Fine-tuning the weights of the attention layers is sufficient to adapt vision transformers to a higher resolution and to other classification tasks.
-
3 Apr 2020 8 repositories listed Syntology ran 2 of 29 samples · 27 unverified · 1 pointer-only (licence)It has been shown that using the first and second order statistics (e.
-
26 Jan 2017 8 repositories listedWe verify the proposed method on a practical problem: person re-identification (re-ID).
-
20 Mar 2020 6 repositories listedTherefore, our multi-branch and multi-scale learning network(MMAL-Net) has good classification ability and robustness for images of different scales.
-
22 Apr 2021 5 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedImageNet-1K serves as the primary dataset for pretraining deep learning models for computer vision tasks.
-
11 Feb 2021 5 repositories listed Syntology ran 8 of 10 samples · 2 unverified · 9 pointer-only (licence)In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset.
-
8 Mar 2020 5 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn this work, we propose a novel framework for fine-grained visual classification to tackle these problems.
-
28 Mar 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models.
-
29 Apr 2021 4 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedOn semi-supervised learning benchmarks we improve performance significantly when only 1% ImageNet labels are available, from 53.
-
12 Jun 2019 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedAppearance information alone is often not sufficient to accurately differentiate between fine-grained visual categories.
-
26 Jan 2019 4 repositories listedSpecifically, for each training image, we first generate attention maps to represent the object's discriminative parts by weakly supervised learning.
-
4 Dec 2017 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Towards addressing this problem, we propose an iterative matrix square root normalization method for fast end-to-end training of global covariance pooling networks.
-
29 Apr 2015 4 repositories listedWe then present a systematic analysis of these networks and show that (1) the bilinear features are highly redundant and can be reduced by an order of magnitude in size without significant loss in accuracy, (2) are also…
-
28 Mar 2024 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedThis paper revives Densely Connected Convolutional Networks (DenseNets) and reveals the underrated effectiveness over predominant ResNet-style architectures.
-
7 Jul 2020 3 repositories listedTraditional learning with ImageNet pre-trained initial weights and SpinalNet classification layers provided the SOTA performance on STL-10, Fruits 360, Bird225, and Caltech-101 datasets.
-
31 Mar 2020 3 repositories listedThe former class can leverage fine-grained semantic relations between data points, but slows convergence in general due to its high training complexity.
-
30 Mar 2020 3 repositories listedIn this work, we introduce a series of architecture modifications that aim to boost neural networks' accuracy, while retaining their GPU training and inference efficiency.
-
11 Feb 2020 3 repositories listed Syntology ran 2 of 5 samples · 3 unverifiedThe proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component.
Syntology lines on 23 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections