Papers › Masked Image Residual Learning for Scaling Deeper Vision Transformers

Masked Image Residual Learning for Scaling Deeper Vision Transformers

25 Sep 2023NeurIPS 2023 11arXiv:2309.14136archive 2025-07-28

Guoxi Huang, Hongtao Fu, Adrian G. Bors

Deeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training. To ease the training of deeper ViTs, we introduce a self-supervised learning framework called Masked Image Residual Learning (MIRL), which significantly alleviates the degradation problem, making scaling ViT along depth a promising direction for performance upgrade. We reformulate the pre-training objective for deeper layers of ViT as learning to recover the residual of the masked image. We provide extensive empirical evidence showing that deeper ViTs can be effectively optimized using MIRL and easily gain accuracy from increased depth. With the same level of computational complexity as ViT-Base and ViT-Large, we instantiate 4.5× and 2× deeper ViTs, dubbed ViT-S-54 and ViT-B-48. The deeper ViT-S-54, costing 3× less than ViT-Large, achieves performance on par with ViT-Large. ViT-B-48 achieves 86.2% top-1 accuracy on ImageNet. On one hand, deeper ViTs pre-trained with MIRL exhibit excellent generalization capabilities on downstream tasks, such as object detection and semantic segmentation. On the other hand, MIRL demonstrates high pre-training efficiency. With less pre-training time, MIRL yields competitive performance compared to other approaches.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

russellllaputa/MIRL officialmentioned in paperpaddlenot reachable when probed 2026-09-17 — repositories for recent papers often appear after camera-ready report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationObject DetectionSelf-Supervised Image ClassificationSelf-Supervised LearningSemantic Segmentationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ImageNet MIRL (ViT-B-48) GFLOPs 67.0 #170 of 1060 Archive leaderboard report
Image Classification ImageNet MIRL (ViT-B-48) Number of params 341M #170 of 1060 Archive leaderboard report
Image Classification ImageNet MIRL (ViT-B-48) Top 1 Accuracy 86.2% #170 of 1060 Archive leaderboard report
Image Classification ImageNet MIRL(ViT-S-54) GFLOPs 18.8 #290 of 1060 Archive leaderboard report
Image Classification ImageNet MIRL(ViT-S-54) Number of params 96M #290 of 1060 Archive leaderboard report
Image Classification ImageNet MIRL(ViT-S-54) Top 1 Accuracy 84.8% #290 of 1060 Archive leaderboard report
Self-Supervised Image Classification ImageNet (finetuned) MIRL (ViT-B-48) Number of Params 341M #16 of 65 Archive leaderboard report
Self-Supervised Image Classification ImageNet (finetuned) MIRL (ViT-B-48) Top 1 Accuracy 86.2% #16 of 65 Archive leaderboard report
Self-Supervised Image Classification ImageNet (finetuned) MIRL (ViT-S-54) Number of Params 96M #27 of 65 Archive leaderboard report
Self-Supervised Image Classification ImageNet (finetuned) MIRL (ViT-S-54) Top 1 Accuracy 84.8% #27 of 65 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections