Papers › Characterizing and Improving the Robustness of Self-Supervised Learning through...

Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations

23 Mar 2021arXiv:2103.12719archive 2025-07-28

Chaitanya K. Ryali, David J. Schwab, Ari S. Morcos

Recent progress in self-supervised learning has demonstrated promising results in multiple visual tasks. An important ingredient in high-performing self-supervised methods is the use of data augmentation by training models to place different augmented views of the same image nearby in embedding space. However, commonly used augmentation pipelines treat images holistically, ignoring the semantic relevance of parts of an image-e.g. a subject vs. a background-which can lead to the learning of spurious correlations. Our work addresses this problem by investigating a class of simple, yet highly effective "background augmentations", which encourage models to focus on semantically-relevant content by discouraging them from focusing on image backgrounds. Through a systematic investigation, we show that background augmentations lead to substantial improvements in performance across a spectrum of state-of-the-art self-supervised methods (MoCo-v2, BYOL, SwAV) on a variety of tasks, e.g. ∼+1-2% gains on ImageNet, enabling performance on par with the supervised baseline. Further, we find the improvement in limited-labels settings is even larger (up to 4.2%). Background augmentations also improve robustness to a number of distribution shifts, including natural adversarial examples, ImageNet-9, adversarial attacks, ImageNet-Renditions. We also make progress in completely unsupervised saliency detection, in the process of generating saliency masks used for background augmentations.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningData AugmentationImage ClassificationRepresentation LearningSaliency DetectionSelf-Supervised LearningUnsupervised Saliency Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ObjectNet BYOL (BG_RM) Top-1 Accuracy 23.9 #83 of 106 Archive leaderboard report
Image Classification ObjectNet SwAV (BG_RM) Top-1 Accuracy 21.9 #86 of 106 Archive leaderboard report
Image Classification ObjectNet MoCo-v2 (BG_Swaps) Top-1 Accuracy 20.8 #88 of 106 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BYOL

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections