Papers › Fine-Grained Visual Classification via Internal Ensemble Learning Transformer

Fine-Grained Visual Classification via Internal Ensemble Learning Transformer

13 Feb 2023IEEE Transactions on Multimedia 2023 2archive 2025-07-28

Qin Xu, Jiahui Wang, Bo Jiang, Bin Luo

Recently, vision transformers (ViTs) have been investigated in fine-grained visual recognition (FGVC) and are now considered state of the art. However, most ViT-based works ignore the different learning performances of the heads in the multihead self-attention (MHSA) mechanism and its layers. To address these issues, in this paper, we propose a novel internal ensemble learning transformer (IELT) for FGVC. The proposed IELT involves three main modules: multi-head voting (MHV) module, cross-layer refinement (CLR) module, and dynamic selection (DS) module. To solve the problem of the inconsistent performances of multiple heads, we propose the MHV module, which considers all of the heads in each layer as weak learners and votes for tokens of discriminative regions as cross-layer feature based on the attention maps and spatial relationships. To effectively mine the cross-layer feature and suppress the noise, the CLR module is proposed, where the refined feature is extracted and the assist logits operation is developed for the final prediction. In addition, a newly designed DS module adjusts the token selection number at each layer by weighting their contributions of the refined feature. In this way, the idea of ensemble learning is combined with the ViT to improve fine-grained feature representation. The experiments demonstrate that our method achieves competitive results compared with the state of the art on five popular FGVC datasets. Source code has been released and can be found at https://github.com/mobulan/IELT.

PaperPDFCode

Code

mobulan/ielt officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationEnsemble LearningFine-Grained Image ClassificationFine-Grained Visual Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification NABirds IELT Accuracy 90.8% #17 of 30 Archive leaderboard report
Fine-Grained Image Classification Oxford 102 Flowers IELT Accuracy 99.64% #1 of 25 Archive leaderboard report
Fine-Grained Image Classification Oxford-IIIT Pet Dataset IELT Accuracy 95.28% #7 of 15 Archive leaderboard report
Fine-Grained Image Classification Stanford Dogs IELT Accuracy 91.8% #12 of 24 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections