Home › Datasets › task › Visual Question Answering (VQA)

Visual Question Answering (VQA) datasets

archive 2025-07-28

144 datasets carry the task tag "Visual Question Answering (VQA)" (the task itself: Visual Question Answering (VQA)), ordered by the archive's paper count. Page 1 of 3: 48 shown of 144. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Visual Question Answering (VQA) datasets 1–48 of 144

The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
Visual Genome contains Visual Question Answering data in a multi-choice setting.
1,256 papers · 15 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images.
366 papers · 7 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
ScienceQA (Science Question Answering)
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
VizWiz (VizWiz-VQA)
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question.
260 papers · 7 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
AI2D (AI2 Diagrams)
AI2 Diagrams (AI2D) is a dataset of over 5000 grade school science diagrams with over 150000 rich annotations, their ground truth syntactic parses, and more than 15000 corresponding multiple choice questions.
207 papers · 1 benchmark
VCR (Visual Commonsense Reasoning)
Visual Commonsense Reasoning (VCR) is a large-scale dataset for cognition-level visual understanding.
179 papers · 13 benchmarks
VisDial (Visual Dialog)
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
A-OKVQA is crowdsourced visual question answering dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer.
154 papers · 1 benchmark
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
MVBench is a comprehensive Multi-modal Video understanding Benchmark.
139 papers · 3 benchmarks
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks.
117 papers · 2 benchmarks
EgoSchema is very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems.
112 papers · 3 benchmarks
Visual7W is a large-scale visual question answering (QA) dataset, with object-level groundings and multimodal answers.
112 papers · 1 benchmark
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
The ReferIt dataset contains 130,525 expressions for referring to 96,654 objects in 19,894 images of natural scenes.
80 papers · 0 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset.
66 papers · 5 benchmarks
COCO-QA is a dataset for visual question answering.
62 papers · 0 benchmarks
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
PMC-VQA is a large-scale medical visual question-answering dataset that contains 227k VQA pairs of 149k images that cover various modalities or diseases.
55 papers · 2 benchmarks
DAQUAR (DAtaset for QUestion Answering on Real-world images) is a dataset of human question answer pairs about images.
54 papers · 0 benchmarks
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
InfographicVQA is a dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations.
52 papers · 1 benchmark
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
TQA (Textbook Question Answering)
The TextbookQuestionAnswering (TQA) dataset is drawn from middle school science curricula.
48 papers · 1 benchmark
WebQA, is a new benchmark for multimodal multihop reasoning in which systems are presented with the same style of data as humans when searching the web: Snippets and Images.
43 papers · 0 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
TDIUC (Task Directed Image Understanding Challenge)
Task Directed Image Understanding Challenge (TDIUC) dataset is a Visual Question Answering dataset which consists of 1.6M questions and 170K images sourced from MS COCO and the Visual Genome Dataset.
39 papers · 1 benchmark
CV-Bench (Cambrian Vision-Centric Benchmark)
The Cambrian Vision-Centric Benchmark (CV-Bench) is designed to address the limitations of existing vision-centric benchmarks by providing a comprehensive evaluation framework for multimodal large language models (MLLMs).
37 papers · 0 benchmarks
InfoSeek (Visual Information Seeking)
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.