Papers › Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

22 Aug 2022arXiv:2208.10442archive 2025-07-28

Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, Furu Wei

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

microsoft/unilm officialpytorch report
lyan62/data-curation mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AllCross-Modal RetrievalImage CaptioningImage ClassificationInstance SegmentationLanguage ModelingLanguage ModellingMasked Language ModelingObject DetectionQuestion AnsweringRetrievalSemantic SegmentationVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningZero-Shot Cross-Modal Retrievalimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval COCO 2014 BEiT-3 Image-to-text R@1 84.8 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 BEiT-3 Image-to-text R@10 98.3 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 BEiT-3 Image-to-text R@5 96.5 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 BEiT-3 Text-to-image R@1 67.2 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 BEiT-3 Text-to-image R@10 87.7 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 BEiT-3 Text-to-image R@5 92.8 #3 of 36 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@1 98.0 #3 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@10 100.0 #3 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@5 100.0 #3 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@1 90.3 #3 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@10 99.5 #3 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@5 98.7 #3 of 27 Archive leaderboard report
Instance Segmentation COCO test-dev BEiT-3 mask AP 54.8 #6 of 112 Archive leaderboard report
Object Detection COCO test-dev BEiT-3 box mAP 63.7 #14 of 225 Archive leaderboard report
Semantic Segmentation ADE20K BEiT-3 Params (M) 1900 #5 of 235 Archive leaderboard report
Semantic Segmentation ADE20K BEiT-3 Validation mIoU 62.8 #5 of 235 Archive leaderboard report
Semantic Segmentation ADE20K val BEiT-3 mIoU 62.8 #1 of 95 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-dev BEiT-3 Accuracy 84.19 #2 of 56 Archive leaderboard report
Visual Question Answering (VQA) VQA v2 test-std BEiT-3 overall 84.03 #1 of 38 Archive leaderboard report
Visual Reasoning NLVR2 Dev BEiT-3 Accuracy 91.51 #1 of 15 Archive leaderboard report
Visual Reasoning NLVR2 Test BEiT-3 Accuracy 92.58 #1 of 14 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@1 94.9 #2 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@10 100.0 #2 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Image-to-text R@5 99.9 #2 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@1 81.5 #2 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@10 97.8 #2 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k BEiT-3 Text-to-image R@5 95.6 #2 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections