Datasets › Conceptual Captions

Conceptual Captions

Introduced by Piyush Sharma et al. in Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning1 Jul 2018 archive 2025-07-28

Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image. Up to this point, the resource most used for this task was the MS-COCO dataset, containing around 120,000 images and 5-way image-caption annotations (produced by paid annotators).

Google's Conceptual Captions dataset has more than 3 million images, paired with natural-language captions. In contrast with the curated style of the MS-COCO images, Conceptual Captions images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. The raw descriptions are harvested from the Alt-text HTML attribute associated with web images. The authors developed an automatic pipeline that extracts, filters, and transforms candidate image/caption pairs, with the goal of achieving a balance of cleanliness, informativeness, fluency, and learnability of the resulting captions.

Source: Conceptual Captions Image Source: Sharma et al

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-to-Image Generation Conceptual Captions Contextual RQ-Transformer FID 9.80 Draft-and-Revise: Effective Image Generation with... — 5 Compare
Image Captioning Conceptual Captions ClipCap (MLP + GPT2 tuning) CIDEr 87.26 ClipCap: CLIP Prefix for Image Captioning rmokady/clip_prefix_caption +3 2 Compare

Papers archive 2025-07-28

6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 352. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Draft-and-Revise: Effective Image Generation with Contextual RQ-Transformer 0 1 9 Jun 2022 not harvested
Autoregressive Image Generation using Residual Quantization 4 1 3 Mar 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
High-Resolution Image Synthesis with Latent Diffusion Models 41 1 20 Dec 2021 ran 19 of 28 samples (9 unverified; 5 pointer-only for licence)
ClipCap: CLIP Prefix for Image Captioning 4 2 18 Nov 2021 ran 3 of 6 samples (3 unverified)
ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis 1 1 19 Aug 2021 not harvested
Taming Transformers for High-Resolution Image Synthesis 13 1 17 Dec 2020 ran 6 of 6 samples (0 unverified; 4 pointer-only for licence)

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Conceptual Captions

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections