Papers › Fully Attentional Networks with Self-emerging Token Labeling

Fully Attentional Networks with Self-emerging Token Labeling

8 Jan 2024ICCV 2023 1arXiv:2401.03844archive 2025-07-28

Bingyin Zhao, Zhiding Yu, Shiyi Lan, Yutao Cheng, Anima Anandkumar, Yingjie Lao, Jose M. Alvarez

Recent studies indicate that Vision Transformers (ViTs) are robust against out-of-distribution scenarios. In particular, the Fully Attentional Network (FAN) - a family of ViT backbones, has achieved state-of-the-art robustness. In this paper, we revisit the FAN models and improve their pre-training with a self-emerging token labeling (STL) framework. Our method contains a two-stage training framework. Specifically, we first train a FAN token labeler (FAN-TL) to generate semantically meaningful patch token labels, followed by a FAN student model training stage that uses both the token labels and the original class label. With the proposed STL framework, our best model based on FAN-L-Hybrid (77.3M parameters) achieves 84.8% Top-1 accuracy and 42.1% mCE on ImageNet-1K and ImageNet-C, and sets a new state-of-the-art for ImageNet-A (46.1%) and ImageNet-R (56.6%) without using extra data, outperforming the original FAN counterpart by significant margins. The proposed framework also demonstrates significantly enhanced performance on downstream tasks such as semantic segmentation, with up to 1.7% improvement in robustness over the counterpart model. Code is available at https://github.com/NVlabs/STL.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

NVlabs/STL officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Domain Generalization ImageNet-A FAN-L-Hybrid+STL Top-1 accuracy % 46.1 #22 of 39 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid+STL Number of params 77M #16 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid+STL Top 1 Accuracy 69.2 #16 of 47 Archive leaderboard report
Domain Generalization ImageNet-C FAN-L-Hybrid+STL mean Corruption Error (mCE) 42.1 #16 of 47 Archive leaderboard report
Domain Generalization ImageNet-R FAN-L-Hybrid+STL Top-1 Error Rate 43.4 #18 of 39 Archive leaderboard report
Semantic Segmentation Cityscapes val FAN-L-Hybrid+STL mIoU 82.8 #31 of 99 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections