Home › Datasets › task › Visual Question Answering (VQA)

Visual Question Answering (VQA) datasets

archive 2025-07-28

144 datasets carry the task tag "Visual Question Answering (VQA)" (the task itself: Visual Question Answering (VQA)), ordered by the archive's paper count. Page 3 of 3: 48 shown of 144. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Visual Question Answering (VQA) datasets 97–144 of 144

Contains 25,165 textual news articles collected from hundreds of news media sites (e.g., Yahoo News, Google News, CNN News.) and 76,516 image posts shared on Flickr social media, which are annotated according to 412 real-world events.
3 papers · 0 benchmarks
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.
2 papers · 0 benchmarks
InstaOrder can be used to understand the geometrical relationships of instances in an image.
2 papers · 0 benchmarks
PoseScript is a dataset that pairs a few thousand 3D human poses from AMASS with rich human-annotated descriptions of the body parts and their spatial relationships.
2 papers · 0 benchmarks
VQA 360° is a dataset for visual question answering on 360° images containing around 17,000 real-world image-question-answer triplets for a variety of question types.
2 papers · 0 benchmarks
Visual Beliefs is a dataset of abstract scenes to study visual beliefs.
2 papers · 0 benchmarks
VizWiz-Priv (Visual Privacy dataset)
VizWiz-Priv includes 8,862 regions showing private content across 5,537 images taken by blind people.
2 papers · 0 benchmarks
A large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering.
2 papers · 0 benchmarks
3U-VQA (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset)
To tackle the challenge of obtaining out-of-distribution (OOD) data for LVQA models, we introduced a novel dataset named 3U-VQA dataset (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset).
1 paper · 0 benchmarks
AesVQA is a dataset that contains 72168 high-quality images and 324756 pairs of aesthetic questions.
1 paper · 0 benchmarks
We introduce the novel task of multimodal puzzle solving, framed within the context of visual question-answering.
1 paper · 1 benchmark
The task of Visual Question Answering (VQA) has been studied extensively on general-domain real-world images.
1 paper · 1 benchmark
CLEVR-MRT (CLEVR: Mental Rotation Tests)
CLEVR Mental Rotation Tests (CLEVR-MRT) is a new version of the CLEVR dataset.
1 paper · 0 benchmarks
CORE-MM is an Open-ended VQA benchmark dataset specifically designed for MLLMs, with a focus on complex reasoning tasks.
1 paper · 1 benchmark
ChiQA (Chinese VQA)
ChiQA is a dataset designed for visual question answering tasks that not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning.
1 paper · 0 benchmarks
DME VQA dataset (Diabetic Macular Edema VQA dataset)
Medical VQA dataset built from the IDRiD and eOphta datasets.
1 paper · 0 benchmarks
first everyday task dataset featuring COT outputs, diverse task designs, detailed re-plan processes, along with SFT and DPO sub-datasets.
1 paper · 0 benchmarks
IllusionAnimalstest Dataset Characteristics IllusionAnimalstest is a generated dataset based on a synthetic collection of animal images, including 10 animal classes: cat, dog, pigeon, butterfly, elephant, horse, deer, snake, fish, and…
1 paper · 0 benchmarks
IllusionChartest Dataset Characteristics IllusionChartest is a generated dataset containing 3,300 samples of images that feature sequences of 3 to 5 random characters.
1 paper · 0 benchmarks
IllusionFashionMNISTtest Dataset Characteristics IllusionFashionMNISTtest is a generated dataset derived from the FashionMNIST dataset.
1 paper · 0 benchmarks
IllusionMNISTtest Dataset Characteristics IllusionMNISTtest is a generated dataset derived from the MNIST dataset.
1 paper · 0 benchmarks
Manually vAlidated Vq2a Examples fRom Image/Caption datasetS (MAVERICS) is a suite of test-only visual question answering datasets.
1 paper · 0 benchmarks
A new in-context visual question answering dataset encompassing interleaved image and EHR data derived from MIMIC-IV and MIMIC-CXR-JPG databases.
1 paper · 0 benchmarks
MediConfusion is a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical Multimodal Large Language Models (MLLMs) from a vision perspective.
1 paper · 0 benchmarks
NEWSKVQA is a new dataset of 12K news videos spanning across 156 hours with 1M multiple-choice question-answer pairs covering 8263 unique entities.
1 paper · 0 benchmarks
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English.
1 paper · 0 benchmarks
Synthetic datasets have successfully been used to probe visual question-answering datasets for their reasoning abilities.
1 paper · 1 benchmark
Super-CLEVR-3D is a visual question answering (VQA) dataset where the questions are about the explicit 3D configuration of the objects from images (i.e.
1 paper · 0 benchmarks
T2 Guiding is a dataset of 1000 images, each with six image labels.
1 paper · 0 benchmarks
Text present in images are not merely strings, they provide useful cues about the image.
1 paper · 0 benchmarks
TinySocial is a dataset to enable research on Social Visual Question Answering.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VQA-MHUG is a 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.
1 paper · 0 benchmarks
VQA-OV (Visual Quality Assessment of Omnidirectional Video)
Collects 60 reference sequences and 540 impaired sequences.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
Visual Choice of Plausible Alternatives (VCOPA) is an evaluation dataset containing 380 VCOPA questions and over 1K images with various topics, which is amenable to automatic evaluation, and present the performance of baseline reasoning…
1 paper · 0 benchmarks
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
WebLI (Web Language Image)
WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering,…
1 paper · 0 benchmarks
A dataset automatically generated using question generation neural models and alt-text video captions from the WebVid dataset, with 3M video-question-answer triplets.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The simply-CLEVR dataset aims to provide a benchmark dataset that can be used for transparent quantitative evaluation of explanation methods (aka heatmaps/XAI methods).
1 paper · 0 benchmarks
uBench (MicroBench)
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.