Papers › Understanding The Robustness in Vision Transformers
Understanding The Robustness in Vision Transformers
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Anima Anandkumar, Jiashi Feng, Jose M. Alvarez
Recent studies show that Vision Transformers(ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code is available at: https://github.com/NVlabs/FAN.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Domain Generalization | ImageNet-A | FAN-Hybrid-L(IN-21K, 384) | Top-1 accuracy % | 74.5 | #7 of 39 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid (IN-22k) | Number of params | 77M | #8 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid (IN-22k) | Top 1 Accuracy | 73.6 | #8 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid (IN-22k) | mean Corruption Error (mCE) | 35.8 | #8 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-B-Hybrid (IN-22k) | Number of params | 50M | #14 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-B-Hybrid (IN-22k) | Top 1 Accuracy | 70.5 | #14 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-B-Hybrid (IN-22k) | mean Corruption Error (mCE) | 41.0 | #14 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid | Number of params | 77M | #20 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid | Top 1 Accuracy | 67.7 | #20 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-C | FAN-L-Hybrid | mean Corruption Error (mCE) | 43.0 | #20 of 47 | Archive leaderboard | report |
| Domain Generalization | ImageNet-R | FAN-Hybrid-L(IN-21K, 384)) | Top-1 Error Rate | 28.9 | #4 of 39 | Archive leaderboard | report |
| Image Classification | ImageNet | FAN-L-Hybrid++ | Number of params | 76.8M | #105 of 1060 | Archive leaderboard | report |
| Image Classification | ImageNet | FAN-L-Hybrid++ | Top 1 Accuracy | 87.1% | #105 of 1060 | Archive leaderboard | report |
| Object Detection | COCO minival | FAN-L-Hybrid | box AP | 55.1 | #51 of 220 | Archive leaderboard | report |
| Semantic Segmentation | Cityscapes val | FAN-L-Hybrid | mIoU | 82.3 | #36 of 99 | Archive leaderboard | report |
| Semantic Segmentation | DensePASS | FAN (MiT-B1) | mIoU | 42.54% | #11 of 36 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections