Papers › Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained...

Penalizing the Hard Example But Not Too Much: A Strong Baseline for Fine-Grained Visual Classification

21 Nov 2022IEEE Transactions on Neural Networks and Learning Systems 2022 11archive 2025-07-28

Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, Yi Yang

Though significant progress has been achieved on fine-grained visual classification (FGVC), severe overfitting still hinders model generalization. A recent study shows that hard samples in the training set can be easily fit, but most existing FGVC methods fail to classify some hard examples in the test set. The reason is that the model overfits those hard examples in the training set, but does not learn to generalize to unseen examples in the test set. In this article, we propose a moderate hard example modulation (MHEM) strategy to properly modulate the hard examples. MHEM encourages the model to not overfit hard examples and offers better generalization and discrimination. First, we introduce three conditions and formulate a general form of a modulated loss function. Second, we instantiate the loss function and provide a strong baseline for FGVC, where the performance of a naive backbone can be boosted and be comparable with recent methods. Moreover, we demonstrate that our baseline can be readily incorporated into the existing methods and empower these methods to be more discriminative. Equipped with our strong baseline, we achieve consistent improvements on three typical FGVC datasets, i.e., CUB-200-2011, Stanford Cars, and FGVC-Aircraft. We hope the idea of moderate hard example modulation will inspire future research work toward more effective fine-grained visual recognition.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Fine-Grained Image ClassificationFine-Grained Visual Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification FGVC Aircraft MHEM (strong ResNet50 baseline) Accuracy 92.9% #34 of 57 Archive leaderboard report
Fine-Grained Image Classification Stanford Cars MHEM (strong ResNet50 baseline) Accuracy 94.2% #50 of 83 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Testfail

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections