Papers › ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder
ViT-NeT: Interpretable Vision Transformers with Neural Tree Decoder
Sangwon Kim; Jaeyeal Nam; Byoung Chul Ko
Vision transformers (ViTs), which have demonstrated a state-of-the-art performance in image classification, can also visualize global interpretations through attention-based contributions. How- ever, the complexity of the model makes it difficult to interpret the decision-making process, and the ambiguity of the attention maps can cause incorrect correlations between image patches. In this study, we propose a new ViT neural tree decoder (ViT-NeT). A ViT acts as a backbone, and to solve its limitations, the output contex- tual image patches are applied to the proposed NeT. The NeT aims to accurately classify fine-grained objects with similar inter-class correlations and different intra-class correlations. In addition, it describes the decision-making process through a tree structure and prototype and en- ables a visual interpretation of the results. The proposed ViT-NeT is designed to not only improve the classification performance but also provide a human-friendly interpretation, which is effective in resolving the trade-off between performance and interpretability. We compared the performance of ViT-NeT with other state-of-art methods using widely used fine-grained visual categorization benchmark datasets and experimentally proved that the proposed method is superior in terms of the classification performance and interpretability. The code and models are publicly available at https://github.com/jumpsnack/ViT-NeT.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Fine-Grained Image Classification | CUB-200-2011 | ViT-NeT | Accuracy | 91.7 | #9 of 30 | Archive leaderboard | report |
| Fine-Grained Image Classification | NABirds | ViT-NeT (SwinV2-B) | Accuracy | 92.5% | #5 of 30 | Archive leaderboard | report |
| Fine-Grained Image Classification | Stanford Cars | ViT-NeT (SwinV2-B) | Accuracy | 95.0% | #25 of 83 | Archive leaderboard | report |
| Fine-Grained Image Classification | Stanford Dogs | ViT-NeT (DeiT-III-B) | Accuracy | 93.6% | #3 of 24 | Archive leaderboard | report |
| Image Classification | iNaturalist | ViT-NeT (SwinV2-B) | Top 1 Accuracy | 81.2 | #6 of 19 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections