Papers › Understanding The Robustness in Vision Transformers

Understanding The Robustness in Vision Transformers

26 Apr 2022arXiv:2204.12451archive 2025-07-28

Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Anima Anandkumar, Jiashi Feng, Jose M. Alvarez

Recent studies show that Vision Transformers(ViTs) exhibit strong robustness against various corruptions. Although this property is partly attributed to the self-attention mechanism, there is still a lack of systematic understanding. In this paper, we examine the role of self-attention in learning robust representations. Our study is motivated by the intriguing properties of the emerging visual grouping in Vision Transformers, which indicates that self-attention may promote robustness through improved mid-level representations. We further propose a family of fully attentional networks (FANs) that strengthen this capability by incorporating an attentional channel processing design. We validate the design comprehensively on various hierarchical backbones. Our model achieves a state-of-the-art 87.1% accuracy and 35.8% mCE on ImageNet-1k and ImageNet-C with 76.8M parameters. We also demonstrate state-of-the-art accuracy and robustness in two downstream tasks: semantic segmentation and object detection. Code is available at: https://github.com/NVlabs/FAN.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

nvlabs/fan officialmentioned in papermentioned on GitHubpytorchNOASSERTION report
NVlabs/STL mentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Domain GeneralizationImage ClassificationObject DetectionSemantic Segmentationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Domain Generalization ImageNet-A FAN-Hybrid-L(IN-21K, 384) Top-1 accuracy % 74.5 #7 of 39 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid (IN-22k) Number of params 77M #8 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid (IN-22k) Top 1 Accuracy 73.6 #8 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid (IN-22k) mean Corruption Error (mCE) 35.8 #8 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-B-Hybrid (IN-22k) Number of params 50M #14 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-B-Hybrid (IN-22k) Top 1 Accuracy 70.5 #14 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-B-Hybrid (IN-22k) mean Corruption Error (mCE) 41.0 #14 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid Number of params 77M #20 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid Top 1 Accuracy 67.7 #20 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid mean Corruption Error (mCE) 43.0 #20 of 47 Archive leaderboard report
Domain Generalization ImageNet-R FAN-Hybrid-L(IN-21K, 384)) Top-1 Error Rate 28.9 #4 of 39 Archive leaderboard report
Image Classification ImageNet FAN-L-Hybrid++ Number of params 76.8M #105 of 1060 Archive leaderboard report
Image Classification ImageNet FAN-L-Hybrid++ Top 1 Accuracy 87.1% #105 of 1060 Archive leaderboard report
Object Detection COCO minival FAN-L-Hybrid box AP 55.1 #51 of 220 Archive leaderboard report
Semantic Segmentation Cityscapes val FAN-L-Hybrid mIoU 82.3 #36 of 99 Archive leaderboard report
Semantic Segmentation DensePASS FAN (MiT-B1) mIoU 42.54% #11 of 36 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections