Papers › Optimizing Relevance Maps of Vision Transformers Improves Robustness

Optimizing Relevance Maps of Vision Transformers Improves Robustness

2 Jun 2022arXiv:2206.01161archive 2025-07-28

Hila Chefer, Idan Schwartz, Lior Wolf

It has been observed that visual classification models often rely mostly on the image background, neglecting the foreground, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and manipulate it such that the model is focused on the foreground object. This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

hila-chefer/robustvit officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationOut-of-Distribution Generalization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ObjectNet AR-L (Opt Relevance) Top-1 Accuracy 52.0 #24 of 106 Archive leaderboard report
Image Classification ObjectNet AR-L (Opt Relevance) Top-5 Accuracy 73.5 #24 of 106 Archive leaderboard report
Image Classification ObjectNet AR-B (Opt Relevance) Top-1 Accuracy 47.1 #32 of 106 Archive leaderboard report
Image Classification ObjectNet AR-B (Opt Relevance) Top-5 Accuracy 70 #32 of 106 Archive leaderboard report
Image Classification ObjectNet AR-L Top-1 Accuracy 46.5 #36 of 106 Archive leaderboard report
Image Classification ObjectNet AR-L Top-5 Accuracy 68.3 #36 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-L (Opt Relevance) Top-1 Accuracy 43.2 #37 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-L (Opt Relevance) Top-5 Accuracy 65.8 #37 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-B (Opt Relevance) Top-1 Accuracy 42.2 #40 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-B (Opt Relevance) Top-5 Accuracy 65.1 #40 of 106 Archive leaderboard report
Image Classification ObjectNet AR-B Top-1 Accuracy 41.4 #42 of 106 Archive leaderboard report
Image Classification ObjectNet AR-B Top-5 Accuracy 63.7 #42 of 106 Archive leaderboard report
Image Classification ObjectNet AR-S (Opt Relevance) Top-1 Accuracy 39.3 #45 of 106 Archive leaderboard report
Image Classification ObjectNet AR-S (Opt Relevance) Top-5 Accuracy 61.7 #45 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-L Top-1 Accuracy 37.4 #48 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-L Top-5 Accuracy 59.5 #48 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-L (Opt Relevance) Top-1 Accuracy 36.3 #49 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-L (Opt Relevance) Top-5 Accuracy 56.6 #49 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-B Top-1 Accuracy 35.1 #54 of 106 Archive leaderboard report
Image Classification ObjectNet ViT-B Top-5 Accuracy 56.4 #54 of 106 Archive leaderboard report
Image Classification ObjectNet AR-S Top-1 Accuracy 34.3 #56 of 106 Archive leaderboard report
Image Classification ObjectNet AR-S Top-5 Accuracy 55.8 #56 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-S (Opt Relevance) Top-1 Accuracy 31.6 #60 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-S (Opt Relevance) Top-5 Accuracy 53 #60 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-L Top-1 Accuracy 31.4 #62 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-L Top-5 Accuracy 48.5 #62 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-S Top-1 Accuracy 28.3 #76 of 106 Archive leaderboard report
Image Classification ObjectNet DeiT-S Top-5 Accuracy 47.3 #76 of 106 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformerVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections