Home › Datasets › task › Image Captioning

Image Captioning datasets

archive 2025-07-28

79 datasets carry the task tag "Image Captioning" (the task itself: Image Captioning), ordered by the archive's paper count. Page 2 of 2: 31 shown of 79. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Image Captioning datasets 49–79 of 79

LAION-COCO is the world’s largest dataset of 600M generated high-quality captions for publicly available web-images.
3 papers · 1 benchmark
OpenCHAIR is a benchmark for evaluating open-vocabulary hallucinations in image captioning models.
3 papers · 0 benchmarks
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
This is an open-source image captions dataset for the aesthetic evaluation of images.
2 papers · 0 benchmarks
DeCOCO is a bilingual (English-German) corpus of image descriptions, where the English part is extracted from the COCO dataset, and the German part are translations by a native German speaker.
2 papers · 0 benchmarks
Egoshots is a 2-month Ego-vision Dataset with Autographer Wearable Camera annotated "for free" with transfer learning.
2 papers · 0 benchmarks
PoseScript is a dataset that pairs a few thousand 3D human poses from AMASS with rich human-annotated descriptions of the body parts and their spatial relationships.
2 papers · 0 benchmarks
VizWiz-Priv (Visual Privacy dataset)
VizWiz-Priv includes 8,862 regions showing private content across 5,537 images taken by blind people.
2 papers · 0 benchmarks
A large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering.
2 papers · 0 benchmarks
3U-VQA (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset)
To tackle the challenge of obtaining out-of-distribution (OOD) data for LVQA models, we introduced a novel dataset named 3U-VQA dataset (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset).
1 paper · 0 benchmarks
This is a dataset for Bengali Captioning from Images.
1 paper · 0 benchmarks
Consists of eye movements and verbal descriptions recorded synchronously over images.
1 paper · 0 benchmarks
ESP (Evaluation for Styled Prompt)
ESP dataset (Evaluation for Styled Prompt dataset) is a benchmark for zero-shot domain-conditional caption generation.
1 paper · 0 benchmarks
GroundCap is a novel grounded image captioning dataset derived from MovieNet, containing 52,350 movie frames with detailed grounded captions.
1 paper · 0 benchmarks
Image Caption Quality Dataset is a dataset of crowdsourced ratings for machine-generated image captions.
1 paper · 0 benchmarks
InFashAI (Inclusive Fashion AI)
AI algorithms, and in particular Machine Learning (ML) algorithms, learn from data tasks that have been traditionally done by humans such as: image classification, facial recognition, linguistic translation etc.
1 paper · 0 benchmarks
A large-scale machine comprehension dataset (based on the COCO images and captions).
1 paper · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English.
1 paper · 0 benchmarks
RPCD (Reddit Photo Critique Dataset)
The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques.
1 paper · 0 benchmarks
SMR IU X-Ray (Simplified Medical Reports)
This paper introduces CPIR-MR (Chained Prompting for Improved Readability of Medical Reports), a method designed to simplify complex chest X-ray reports for better patient understanding.
1 paper · 0 benchmarks
T2 Guiding is a dataset of 1000 images, each with six image labels.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
WebLI (Web Language Image)
WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering,…
1 paper · 0 benchmarks
WikiWeb2M (Wikipedia Webpage 2M)
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A Filipino multi-modal language dataset for text+visual tasks.
0 papers · 0 benchmarks
ESP Dataset (Evaluation for Styled Prompt datase)
ESP dataset (Evaluation for Styled Prompt dataset) is a new benchmark for zero-shot domain-conditional caption generation.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.