Papers › Scaling Up Vision-Language Pre-training for Image Captioning

Scaling Up Vision-Language Pre-training for Image Captioning

24 Nov 2021CVPR 2022 1arXiv:2111.12233archive 2025-07-28

Xiaowei Hu, Zhe Gan, JianFeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, Lijuan Wang

In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art VinVL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million image-text pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail visual concepts when used in a zero-shot manner.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeImage Captioning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Captioning COCO Captions LEMON BLEU-4 42.6 #7 of 41 Archive leaderboard report
Image Captioning COCO Captions LEMON CIDER 145.5 #7 of 41 Archive leaderboard report
Image Captioning COCO Captions LEMON METEOR 31.4 #7 of 41 Archive leaderboard report
Image Captioning COCO Captions LEMON SPICE 25.5 #7 of 41 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team B1 85.62 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team B2 71.36 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team B3 53.62 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team B4 34.65 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team CIDEr 114.25 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team METEOR 31.27 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team ROUGE-L 61.2 #3 of 12 Archive leaderboard report
Image Captioning nocaps-XD entire Microsoft Cognitive Services team SPICE 14.85 #3 of 12 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_large CIDEr 116.9 #4 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_large Pre-train (#images) 200M #4 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_large SPICE 15.8 #4 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_base CIDEr 107.7 #8 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_base Pre-train (#images) 200M #8 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain LEMON_base SPICE 14.7 #8 of 11 Archive leaderboard report
Image Captioning nocaps-val-near-domain LEMON_large CIDEr 113.3 #4 of 10 Archive leaderboard report
Image Captioning nocaps-val-near-domain LEMON_large Pre-train (#images) 200M #4 of 10 Archive leaderboard report
Image Captioning nocaps-val-near-domain LEMON_large SPICE 15.1 #4 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain LEMON_large CIDEr 111.3 #7 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain LEMON_large Pretrain (#images) 200M #7 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain LEMON_large SPICE 14.0 #7 of 10 Archive leaderboard report
Image Captioning nocaps-val-overall LEMON_large CIDEr 113.4 #4 of 11 Archive leaderboard report
Image Captioning nocaps-val-overall LEMON_large Pretrain (#images) 200M #4 of 11 Archive leaderboard report
Image Captioning nocaps-val-overall LEMON_large SPICE 15.0 #4 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections