Papers › Multi-branch and Multi-scale Attention Learning for Fine-Grained Visual Categorization

Multi-branch and Multi-scale Attention Learning for Fine-Grained Visual Categorization

20 Mar 2020arXiv:2003.09150archive 2025-07-28

Fan Zhang, Meng Li, Guisheng Zhai, Yizhao Liu

ImageNet Large Scale Visual Recognition Challenge (ILSVRC) is one of the most authoritative academic competitions in the field of Computer Vision (CV) in recent years. But applying ILSVRC's annual champion directly to fine-grained visual categorization (FGVC) tasks does not achieve good performance. To FGVC tasks, the small inter-class variations and the large intra-class variations make it a challenging problem. Our attention object location module (AOLM) can predict the position of the object and attention part proposal module (APPM) can propose informative part regions without the need of bounding-box or part annotations. The obtained object images not only contain almost the entire structure of the object, but also contains more details, part images have many different scales and more fine-grained features, and the raw images contain the complete object. The three kinds of training images are supervised by our multi-branch network. Therefore, our multi-branch and multi-scale learning network(MMAL-Net) has good classification ability and robustness for images of different scales. Our approach can be trained end-to-end, while provides short inference time. Through the comprehensive experiments demonstrate that our approach can achieves state-of-the-art results on CUB-200-2011, FGVC-Aircraft and Stanford Cars datasets. Our code will be available at https://github.com/ZF1044404254/MMAL-Net

PaperPDFCode

Code

ZF1044404254/MMAL-Net officialmentioned in papermentioned on GitHubpytorch report
ZF1044404254/TBMSL-Net officialmentioned on GitHubpytorch report
1170500804/tbmsl mentioned on GitHubpytorch report
ZF4444/MMAL-Net mentioned on GitHubpytorch report
dreamercv/HSResNet-MMAL mentioned on GitHubpytorch report
mv-lab/ViT-FGVC8 mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Fine-Grained Image ClassificationFine-Grained Image RecognitionFine-Grained Visual CategorizationObjectObject Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification CUB-200-2011 TBMSL-Net Accuracy 89.6 #14 of 30 Archive leaderboard report
Fine-Grained Image Classification FGVC Aircraft TBMSL-Net Accuracy 94.7% #5 of 57 Archive leaderboard report
Fine-Grained Image Classification Stanford Cars TBMSL-Net Accuracy 95.0% #23 of 83 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Adaptive NMS

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections