Datasets › COCO Captions

COCO Captions

Introduced by Xinlei Chen et al. in Microsoft COCO Captions: Data Collection and Evaluation Server archive 2025-07-28

COCO Captions contains over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions are be provided for each image.

Source: Microsoft COCO Captions: Data Collection and Evaluation Server

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Image Captioning COCO Captions mPLUG BLEU-4 46.5 mPLUG: Effective and Efficient Vision-Language Learning... modelscope/modelscope +2 41 Compare
Text Generation COCO Captions LeakGAN BLEU-2 0.950 Long Text Generation via Adversarial Training with... CR-Gjx/LeakGAN +5 5 Compare
Image Captioning COCO Captions test From Captions to Visual Concepts and Back BLEU-4 56.7 From Captions to Visual Concepts and Back s-gupta/visual-concepts 2 Compare
Concept-To-Text Generation COCO Captions tecpic BLEU-2 2 Fake News Detection as Natural Language Inference zake7749/WSDM-Cup-2019 1 Compare

Papers archive 2025-07-28

30 shown of 40 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 203. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation? 1 2 16 Apr 2024 not harvested
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 1 29 May 2023 ran 15 of 42 samples (27 unverified)
FuseCap: Leveraging Large Language Models for Enriched Fused Image Captions 1 1 28 May 2023 not harvested
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset 1 1 17 Apr 2023 not harvested
Prismer: A Vision-Language Model with Multi-Task Experts 2 1 4 Mar 2023 not harvested
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 3 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
Position-guided Text Prompt for Vision-Language Pre-training 1 1 19 Dec 2022 ran 3 of 5 samples (2 unverified)
Text-Only Training for Image Captioning using Noise-Injected CLIP 4 1 1 Nov 2022 ran 4 of 5 samples (1 unverified; 1 pointer-only for licence)
SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation 1 1 30 Sep 2022 not harvested
Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning 1 1 13 Aug 2022 not harvested
Prompt Tuning for Generative Multimodal Pretrained Models 1 1 4 Aug 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
GRIT: Faster and Better Image captioning Transformer Using Dual Visual Features 2 1 20 Jul 2022 not harvested
GIT: A Generative Image-to-text Transformer for Vision and Language 1 1 27 May 2022 ran 0 of 21 samples (21 unverified)
Fine-grained Image Captioning with CLIP Reward 1 1 26 May 2022 not harvested
mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections 3 1 24 May 2022 not harvested
Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning 1 3 9 May 2022 not harvested
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 ran 9 of 17 samples (8 unverified)
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework 4 1 7 Feb 2022 ran 1 of 1 samples (0 unverified)
Scaling Up Vision-Language Pre-training for Image Captioning 0 1 24 Nov 2021 not harvested
L-Verse: Bidirectional Generation Between Image and Text 1 1 22 Nov 2021 not harvested
ClipCap: CLIP Prefix for Image Captioning 4 2 18 Nov 2021 ran 3 of 6 samples (3 unverified)
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts 1 1 16 Nov 2021 ran 1 of 1 samples (0 unverified)
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation 0 1 16 Nov 2021 not harvested
RefineCap: Concept-Aware Refinement for Image Captioning 0 1 8 Sep 2021 not harvested
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision 2 1 24 Aug 2021 ran 18 of 37 samples (19 unverified; 28 pointer-only for licence)
VinVL: Revisiting Visual Representations in Vision-Language Models 7 1 2 Jan 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
VirTex: Learning Visual Representations from Textual Annotations 3 1 11 Jun 2020 ran 0 of 1 samples (1 unverified)
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks 4 1 13 Apr 2020 ran 11 of 23 samples (12 unverified)
X-Linear Attention Networks for Image Captioning 2 1 31 Mar 2020 not harvested
A Better Variant of Self-Critical Sequence Training 1 1 22 Mar 2020 not harvested

The full list of 40 is in the JSON twin.

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • COCO Captions
  • COCO Captions test
  • Image Captioning on COCO Captions
  • COCO Captions Karpathy Test

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections