Datasets › Conceptual Captions
Conceptual Captions
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image. Up to this point, the resource most used for this task was the MS-COCO dataset, containing around 120,000 images and 5-way image-caption annotations (produced by paid annotators).
Google's Conceptual Captions dataset has more than 3 million images, paired with natural-language captions. In contrast with the curated style of the MS-COCO images, Conceptual Captions images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. The raw descriptions are harvested from the Alt-text HTML attribute associated with web images. The authors developed an automatic pipeline that extracts, filters, and transforms candidate image/caption pairs, with the goal of achieving a balance of cleanliness, informativeness, fluency, and learnability of the resulting captions.
Source: Conceptual Captions Image Source: Sharma et al
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Text-to-Image Generation | Conceptual Captions | Contextual RQ-Transformer FID 9.80 | Draft-and-Revise: Effective Image Generation with... | — | 5 | Compare |
| Image Captioning | Conceptual Captions | ClipCap (MLP + GPT2 tuning) CIDEr 87.26 | ClipCap: CLIP Prefix for Image Captioning | rmokady/clip_prefix_caption +3 | 2 | Compare |
Papers archive 2025-07-28
6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 352. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Draft-and-Revise: Effective Image Generation with Contextual RQ-Transformer | 0 | 1 | 9 Jun 2022 | not harvested |
| Autoregressive Image Generation using Residual Quantization | 4 | 1 | 3 Mar 2022 | ran 1 of 1 samples (0 unverified; 1 pointer-only for licence) |
| High-Resolution Image Synthesis with Latent Diffusion Models | 41 | 1 | 20 Dec 2021 | ran 19 of 28 samples (9 unverified; 5 pointer-only for licence) |
| ClipCap: CLIP Prefix for Image Captioning | 4 | 2 | 18 Nov 2021 | ran 3 of 6 samples (3 unverified) |
| ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis | 1 | 1 | 19 Aug 2021 | not harvested |
| Taming Transformers for High-Resolution Image Synthesis | 13 | 1 | 17 Dec 2020 | ran 6 of 6 samples (0 unverified; 4 pointer-only for licence) |
Dataset loaders archive 2025-07-28
3 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Conceptual Captions
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections