Browse State-of-the-Art › Fine-Grained Visual Recognition
Fine-Grained Visual Recognition
43 papers with code · 4 benchmarks · 9 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CUB-200-2011 (1 row) | Selfsynthx | Enhancing Cognition and Explainability of Multimodal Foundation... | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| FGVC-Aircraft (1 row) | Selfsynthx | Enhancing Cognition and Explainability of Multimodal Foundation... | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| New Plant Diseases Dataset (1 row) | Selfsynthx | Enhancing Cognition and Explainability of Multimodal Foundation... | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| Stanford Dogs (1 row) | Selfsynthx | Enhancing Cognition and Explainability of Multimodal Foundation... | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
9 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 43 papers with code (72 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Apr 2020 8 repositories listedIn this work we explore the task of instance segmentation with attribute localization, which unifies instance segmentation (detect and segment each object instance) and fine-grained visual attribute categorization…
-
29 Apr 2015 4 repositories listedWe then present a systematic analysis of these networks and show that (1) the bilinear features are highly redundant and can be reduced by an order of magnitude in size without significant loss in accuracy, (2) are also…
-
15 Apr 2019 3 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)The proposed methods are highly modular, readily plugged into existing deep CNNs.
-
11 Jan 2019 3 repositories listedIn this paper, we propose a deep convolutional neural network for learning the embeddings of images in order to capture the notion of visual similarity.
-
20 Mar 2024 2 repositories listed Syntology ran 5 of 7 samples · 2 unverifiedNotably, our approach demonstrates a significant improvement in performance on 5 fine-grained visual recognition benchmarks, 11 few-shot image recognition datasets, and the 2 object detection datasets under the…
-
29 Apr 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedExisting computer vision research in artwork struggles with artwork's fine-grained attributes recognition and lack of curated annotated datasets due to their costly creation.
-
18 Apr 2020 2 repositories listedThis paper introduces a novel dataset FeatherV1, containing 28, 272 images of feathers categorized by 595 bird species.
-
31 Mar 2020 2 repositories listedRecent progress on fine-grained visual recognition and visual question answering has featured Bilinear Pooling, which effectively models the 2ⁿᵈ order interactions across multi-modal inputs.
-
3 Mar 2019 2 repositories listedInspired by the fact that successive CNN layers represent the image with increasing levels of abstraction, we compressed our deep ranking model to a single CNN by coupling activations from multiple intermediate layers…
-
26 Jul 2018 2 repositories listedFine-grained visual recognition is challenging because it highly relies on the modeling of various semantic parts and fine-grained feature learning.
-
18 Nov 2015 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedBeyond classification, we further validate the saliency of the learnt representations via their attribute concentration and hierarchy recovery properties, achieving 10-25% relative gains on the softmax classifier and…
-
19 Feb 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Large multimodal models (LMMs) have shown impressive capabilities in a wide range of visual tasks.
-
25 Jan 2025 1 repository listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks.
-
20 Oct 2024 1 repository listedThis paper presents a novel approach for Fine-Grained Visual Classification (FGVC) by exploring Graph Neural Networks (GNNs) to facilitate high-order feature interactions, with a specific focus on constructing both…
-
3 Sep 2024 1 repository listedMoreover, in the context of building FGVR models with limited data, these irrelevant features can dominate the training process, overshadowing more useful, generalizable discriminative features.
-
3 Sep 2024 1 repository listedSpecifically, we propose two novel methods: Generative Class Prompt Learning (GCPL) and Contrastive Multi-class Prompt Learning (CoMPLe).
-
25 Aug 2024 1 repository listedIn the second stage, we introduce CCFG-Net, a network architecture that integrates classification and contrastive learning in both Euclidean and Angular spaces, in which contrastive learning is applied in both…
-
28 Jun 2024 1 repository listedThe emerging task of fine-grained image classification in low-data regimes assumes the presence of low inter-class variance and large intra-class variation along with a highly limited amount of training samples per…
-
3 Jun 2024 1 repository listedUnderstanding Activities of Daily Living (ADLs) is a crucial step for different applications including assistive robots, smart homes, and healthcare.
-
23 Nov 2023 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedWe explore constructing the class hierarchy into a graph, with its nodes representing the textual or image features of each category.
-
30 Mar 2023 1 repository listedThis leads traditional novel category discovery (NCD) methods to be incapacitated for GCD, due to their assumption of unlabeled data are only from novel categories.
-
3 Mar 2023 1 repository listedSpecifically, we fit the GradCAM with a branch with limited fitting capacity, which allows the branch to capture the common rationales and discard the less common discriminative patterns.
-
13 Feb 2023 1 repository listedThe proposed IELT involves three main modules: multi-head voting (MHV) module, cross-layer refinement (CLR) module, and dynamic selection (DS) module.
-
1 Jan 2023 1 repository listedDespite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2 images.
-
28 Dec 2022 1 repository listedThis framework, namely PArt-guided Relational Transformers (PART), is proposed to learn the discriminative part features with an automatic part discovery module, and to explore the intrinsic correlations with a feature…
-
21 Nov 2022 1 repository listedSecond, we instantiate the loss function and provide a strong baseline for FGVC, where the performance of a naive backbone can be boosted and be comparable with recent methods.
-
1 Aug 2022 1 repository listedIn low data regimes, a network often struggles to choose the correct regions for recognition and tends to overfit spurious correlated patterns from the training data.
-
25 Jul 2022 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedPractical real world datasets with plentiful categories introduce new challenges for unsupervised domain adaptation like small inter-class discriminability, that existing approaches relying on domain invariance alone…
-
26 May 2022 1 repository listed Syntology ran 3 of 17 samples · 14 unverifiedInspired by this observation, we propose a network branch dedicated to magnifying the importance of small eigenvalues.
-
31 Mar 2022 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedCompMap first asks a VL model to generate primitive concept activations with text prompts, and then learns to construct a composition model that maps the primitive concept activations (e.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections