Home › Datasets › task › Image Captioning
Image Captioning datasets
archive 2025-07-28
79 datasets carry the task tag "Image Captioning" (the task itself: Image Captioning), ordered by the archive's paper count. Page 1 of 2: 48 shown of 79. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Image Captioning datasets 1–48 of 79
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The Flickr30k dataset contains 31,000 images collected from Flickr, together with 5 reference sentences provided by human annotators.
880 papers · 9 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question.
260 papers · 7 benchmarks
COCO Captions contains over one and a half million captions describing over 330,000 images.
203 papers · 4 benchmarks
The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
The nocaps benchmark consists of 166,100 human-generated captions describing 15,100 images from the OpenImages validation and test sets.
175 papers · 13 benchmarks
Contains 145k captions for 28k images.
98 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
90 papers · 13 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
The ReferIt dataset contains 130,525 expressions for referring to 96,654 objects in 19,894 images of natural scenes.
80 papers · 0 benchmarks
RSICD (Remote Sensing Image Captioning Dataset)
70 papers · 3 benchmarks
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language.
63 papers · 3 benchmarks
A new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes.
48 papers · 0 benchmarks
A new large-scale dataset for referring expressions, based on MS-COCO.
46 papers · 2 benchmarks
Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits.
43 papers · 4 benchmarks
UT Zappos50K is a large shoe dataset consisting of 50,025 catalog images collected from Zappos.com.
32 papers · 2 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
FlickrStyle10K is collected and built on Flickr30K image caption dataset.
24 papers · 2 benchmarks
WIDER (Web Image Dataset for Event Recognition)
WIDER is a dataset for complex event recognition from static images.
22 papers · 1 benchmark
COCO-CN is a bilingual image description dataset enriching MS-COCO with manually written Chinese sentences and tags.
21 papers · 1 benchmark
IU X-ray (Demner-Fushman et al., 2016) is a set of chest X-ray images paired with their corresponding diagnostic reports.
19 papers · 2 benchmarks
SCICAP is a large-scale image captioning dataset that contains real-world scientific figures and captions.
18 papers · 1 benchmark
Consists of over 39,000 images originating from people who are blind that are each paired with five captions.
12 papers · 0 benchmarks
Object HalBench is a benchmark used to evaluate the performance of Language Models, particularly those that are multimodal (i.e., they can process and generate both text and images).
11 papers · 1 benchmark
A collection that allows researchers to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results.
11 papers · 0 benchmarks
PFN-PIC (PFN Picking Instructions for Commodities Dataset)
This dataset is a collection of spoken language instructions for a robotic system to pick and place common objects.
7 papers · 0 benchmarks
CITE is a crowd-sourced resource for multimodal discourse: this resource characterises inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
6 papers · 1 benchmark
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
Peir Gross (Jing et al., 2018) was collected with descriptions in the Gross sub-collection from PEIR digital library, resulting in 7.442 image-caption pairs from 21 different sub-categories.
6 papers · 1 benchmark
UIT-ViIC contains manually written captions for images from Microsoft COCO dataset relating to sports played with ball.
6 papers · 0 benchmarks
This dataset consists of images and annotations in Bengali.
5 papers · 1 benchmark
Contains 8k flickr Images with captions.
5 papers · 2 benchmarks
PoMo consists of more than 231K sentences with post-modifiers and associated facts extracted from Wikidata for around 57K unique entities.
5 papers · 0 benchmarks
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
Tencent ML-Images is a large open-source multi-label image database, including 17,609,752 training and 88,739 validation image URLs, which are annotated with up to 11,166 categories.
5 papers · 0 benchmarks
Concadia is a publicly available Wikipedia-based corpus, which consists of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.
4 papers · 0 benchmarks
A new language-guided image editing dataset that contains a large number of real image pairs with corresponding editing instructions.
4 papers · 0 benchmarks
The Polaris dataset offers a large-scale, diverse benchmark for evaluating metrics for image captioning, surpassing existing datasets in terms of size, caption diversity, number of human judgments, and granularity of the evaluations.
4 papers · 0 benchmarks
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions.
4 papers · 0 benchmarks
Hephaestus (Hephaestus: A large scale multitask dataset towards InSAR understanding)
Hephaestus is the first large-scale InSAR dataset.
3 papers · 0 benchmarks
Please refer: https://github.com/google/imageinwords/blob/main/datasets/IIW-400/README.md
3 papers · 0 benchmarks
The Image and Video Advertisements collection consists of an image dataset of 64,832 image ads, and a video dataset of 3,477 ads.
3 papers · 0 benchmarks
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.