Papers › ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

22 Jan 2020arXiv:2001.07966archive 2025-07-28

Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, Arun Sacheti

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relationship between them. The model is pre-trained on four tasks simultaneously: Masked Language Modeling (MLM), Masked Object Classification (MOC), Masked Region Feature Regression (MRFR), and Image Text Matching (ITM). To further enhance the pre-training quality, we have collected a Large-scale weAk-supervised Image-Text (LAIT) dataset from Web. We first pre-train the model on this dataset, then conduct a second stage pre-training on Conceptual Captions and SBU Captions. Our experiments show that multi-stage pre-training strategy outperforms single-stage pre-training. We also fine-tune and evaluate our pre-trained ImageBERT model on image retrieval and text retrieval tasks, and achieve new state-of-the-art results on both MSCOCO and Flickr30k datasets.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image RetrievalImage-text matchingLanguage ModelingLanguage ModellingMasked Language ModelingRetrievalText MatchingText RetrievalZero-Shot Cross-Modal Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Image-to-text R@1 44.0 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Image-to-text R@10 80.4 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Image-to-text R@5 71.2 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Text-to-image R@1 32.3 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Text-to-image R@10 70.2 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 ImageBERT Text-to-image R@5 59.0 #17 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Image-to-text R@1 70.7 #20 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Image-to-text R@10 94.0 #20 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Image-to-text R@5 90.2 #20 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Text-to-image R@1 54.3 #20 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Text-to-image R@10 87.5 #20 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k ImageBERT Text-to-image R@5 79.6 #20 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections