Papers › Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail...

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

17 Feb 2021CVPR 2021 1arXiv:2102.08981archive 2025-07-28

Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

google-research-datasets/conceptual-12m officialmentioned in papermentioned on GitHub report
facebookresearch/meru mentioned on GitHubpytorch report
gicheonkang/gst-visdial mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Caption GenerationDiversityImage CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Datasets

Introduced by this paper, per the archive.

CC12M

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Captioning nocaps-val-in-domain Enc-Dec CIDEr 92.6 #11 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain Enc-Dec Pre-train (#images) 15M #11 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain Enc-Dec SPICE 12.5 #11 of 11 Archive leaderboard report
Image Captioning nocaps-val-near-domain Enc-Dec CIDEr 88.3 #10 of 10 Archive leaderboard report
Image Captioning nocaps-val-near-domain Enc-Dec SPICE 12.1 #10 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain Enc-Dec CIDEr 94.5 #9 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain Enc-Dec SPICE 11.9 #9 of 10 Archive leaderboard report
Image Captioning nocaps-val-overall Enc-Dec CIDEr 90.2 #10 of 11 Archive leaderboard report
Image Captioning nocaps-val-overall Enc-Dec SPICE 12.1 #10 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections