Papers › ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder

ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder

17 Jul 2022ICML 2022 7archive 2025-07-28

Sangwon Kim; Jaeyeal Nam; Byoung Chul Ko

Vision transformers (ViTs), which have demonstrated a state-of-the-art performance in image classification, can also visualize global interpretations through attention-based contributions. How- ever, the complexity of the model makes it difficult to interpret the decision-making process, and the ambiguity of the attention maps can cause incorrect correlations between image patches. In this study, we propose a new ViT neural tree decoder (ViT-NeT). A ViT acts as a backbone, and to solve its limitations, the output contex- tual image patches are applied to the proposed NeT. The NeT aims to accurately classify fine-grained objects with similar inter-class correlations and different intra-class correlations. In addition, it describes the decision-making process through a tree structure and prototype and en- ables a visual interpretation of the results. The proposed ViT-NeT is designed to not only improve the classification performance but also provide a human-friendly interpretation, which is effective in resolving the trade-off between performance and interpretability. We compared the performance of ViT-NeT with other state-of-art methods using widely used fine-grained visual categorization benchmark datasets and experimentally proved that the proposed method is superior in terms of the classification performance and interpretability. The code and models are publicly available at https://github.com/jumpsnack/ViT-NeT.

PaperPDFCode

Code

jumpsnack/ViT-NeT officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Decision MakingDecoderFine-Grained Image ClassificationFine-Grained Visual CategorizationImage Classificationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification CUB-200-2011 ViT-NeT Accuracy 91.7 #9 of 30 Archive leaderboard report
Fine-Grained Image Classification NABirds ViT-NeT (SwinV2-B) Accuracy 92.5% #5 of 30 Archive leaderboard report
Fine-Grained Image Classification Stanford Cars ViT-NeT (SwinV2-B) Accuracy 95.0% #25 of 83 Archive leaderboard report
Fine-Grained Image Classification Stanford Dogs ViT-NeT (DeiT-III-B) Accuracy 93.6% #3 of 24 Archive leaderboard report
Image Classification iNaturalist ViT-NeT (SwinV2-B) Top 1 Accuracy 81.2 #6 of 19 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections