Browse State-of-the-Art › Fine-Grained Image Recognition
Fine-Grained Image Recognition
39 papers with code · 4 benchmarks · 12 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| OVEN (5 rows) | PaLI-X | PaLI-X: On Scaling up a Multilingual Vision and Language Model | code | Syntology ran 6 of 7 samples · 1 unverified | Compare |
| CNFOOD-241-Chen (2 rows) | Res-VMamba-S | Res-VMamba: Fine-Grained Food Category Visual Classification Using... | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| CUB-200-2011 (2 rows) | PIM | A Novel Plug-in Module for Fine-Grained Visual Classification | code | — | Compare |
| CUB Birds (2 rows) | HOI-Net | High-Order-Interaction for weakly supervised Fine-Grained Visual... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 39 papers with code (72 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Mar 2020 6 repositories listedTherefore, our multi-branch and multi-scale learning network(MMAL-Net) has good classification ability and robustness for images of different scales.
-
4 Dec 2017 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Towards addressing this problem, we propose an iterative matrix square root normalization method for fast end-to-end training of global covariance pooling networks.
-
1 Oct 2017 3 repositories listedTwo losses are proposed to guide the multi-task learning of channel grouping and part classification, which encourages MA-CNN to generate more discriminative parts from feature channels and learn better fine-grained…
-
14 Oct 2023 2 repositories listedHowever, the absence of a unified open-source software library covering various paradigms in FGIR poses a significant challenge for researchers and practitioners in the field.
-
29 May 2023 2 repositories listed Syntology ran 6 of 7 samples · 1 unverifiedWe present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture.
-
22 Feb 2023 2 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedLarge-scale multi-modal pre-training models such as CLIP and PaLI exhibit strong generalization on various visual domains and tasks.
-
20 Mar 2021 2 repositories listedWe formulate it as a multi-agent reinforcement learning (MARL) problem, where each agent learns an augmentation policy for each patch based on its content together with the semantics of the whole image.
-
1 Jun 2019 2 repositories listedIn this paper, we propose a novel "Destruction and Construction Learning" (DCL) method to enhance the difficulty of fine-grained recognition and exercise the classification model to acquire expert knowledge.
-
31 Dec 2024 1 repository listedUltra-fine-grained image recognition (UFGIR) is a challenging task that involves classifying images within a macro-category.
-
17 Sep 2024 1 repository listedUltra-fine-grained image recognition (UFGIR) categorizes objects with extremely small differences between classes, such as distinguishing between cultivars within the same species, as opposed to species-level…
-
3 Sep 2024 1 repository listedMoreover, in the context of building FGVR models with limited data, these irrelevant features can dominate the training process, overshadowing more useful, generalizable discriminative features.
-
17 Jul 2024 1 repository listedFine-grained recognition involves the classification of images from subordinate macro-categories, and it is challenging due to small inter-class differences.
-
24 Feb 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Our findings elucidate that our proposed methodology establishes a new benchmark for SOTA performance in food recognition on the CNFOOD-241 dataset.
-
1 Sep 2023 1 repository listed Syntology ran 3 of 7 samples · 4 unverified · 7 pointer-only (licence)Since images belonging to the same meta-category usually share similar visual appearances, mining discriminative visual cues is the key to distinguishing fine-grained categories.
-
21 Nov 2022 1 repository listedAlthough the previously proposed method of using both face data without facemasks and those with synthesised facemasks for training improved the recognition performance, a certain drop in the recognition performance for…
-
28 Jun 2022 1 repository listedMolecular and morphological characters, as important parts of biological taxonomy, are contradictory but need to be integrated.
-
24 Mar 2022 1 repository listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)A visual counterfactual explanation replaces image regions in a query image with regions from a distractor image such that the system's decision on the transformed image changes to the distractor class.
-
8 Feb 2022 1 repository listedVisual classification can be divided into coarse-grained and fine-grained classification.
-
13 Nov 2021 1 repository listedOf those, methods based on bilinear pooling are one of the main categories for computing the interaction between deep features and have shown high effectiveness.
-
5 Nov 2021 1 repository listedWe then train a fine-grained textual similarity model that matches image descriptions with documents on a sentence-level basis.
-
23 Oct 2021 1 repository listedWe address this by proposing an end-to-end CNN model, which learns meaningful features linking fine-grained changes using our novel attention mechanism.
-
7 Sep 2021 1 repository listedThis paper contributes a new high-quality dataset for hand gesture recognition in hand hygiene systems, named "MFH".
-
19 Aug 2021 1 repository listed Syntology ran 8 of 11 samples · 3 unverifiedUnlike most existing methods that learn visual attention based on conventional likelihood, we propose to learn the attention with counterfactual causality, which provides a tool to measure the attention quality and a…
-
30 Jun 2021 1 repository listedSelf-supervised contrastive learning has demonstrated great potential in learning visual representations.
-
18 Mar 2021 1 repository listedInterestingly, ViT achieves results superior to CNN baselines with 80.
-
3 Dec 2020 1 repository listedWe propose the Neural Prototype Tree (ProtoTree), an intrinsically interpretable deep learning method for fine-grained image recognition.
-
26 Nov 2020 1 repository listed Syntology ran 4 of 5 samples · 1 unverifiedWe evaluate the transfer performance of 13 top self-supervised models on 40 downstream tasks, including many-shot and few-shot recognition, object detection, and dense prediction.
-
8 Sep 2020 1 repository listedThe evaluation information is backpropagated and forces the classification stream to improve its awareness of visual attention, which helps classification.
-
18 Jun 2020 1 repository listedOne of the difficulties of this competition is how to use unlabeled data.
-
28 May 2020 1 repository listedIdentifying prescription medications is a frequent task for patients and medical professionals; however, this is an error-prone task as many pills have similar appearances (e.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections