Papers › Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail...
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut
The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.
In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Captioning | nocaps-val-in-domain | Enc-Dec | CIDEr | 92.6 | #11 of 11 | Archive leaderboard | report |
| Image Captioning | nocaps-val-in-domain | Enc-Dec | Pre-train (#images) | 15M | #11 of 11 | Archive leaderboard | report |
| Image Captioning | nocaps-val-in-domain | Enc-Dec | SPICE | 12.5 | #11 of 11 | Archive leaderboard | report |
| Image Captioning | nocaps-val-near-domain | Enc-Dec | CIDEr | 88.3 | #10 of 10 | Archive leaderboard | report |
| Image Captioning | nocaps-val-near-domain | Enc-Dec | SPICE | 12.1 | #10 of 10 | Archive leaderboard | report |
| Image Captioning | nocaps-val-out-domain | Enc-Dec | CIDEr | 94.5 | #9 of 10 | Archive leaderboard | report |
| Image Captioning | nocaps-val-out-domain | Enc-Dec | SPICE | 11.9 | #9 of 10 | Archive leaderboard | report |
| Image Captioning | nocaps-val-overall | Enc-Dec | CIDEr | 90.2 | #10 of 11 | Archive leaderboard | report |
| Image Captioning | nocaps-val-overall | Enc-Dec | SPICE | 12.1 | #10 of 11 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections