Home › Datasets › task › Visual Question Answering (VQA)
Visual Question Answering (VQA) datasets
archive 2025-07-28
144 datasets carry the task tag "Visual Question Answering (VQA)" (the task itself: Visual Question Answering (VQA)), ordered by the archive's paper count. Page 2 of 3: 48 shown of 144. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Visual Question Answering (VQA) datasets 49–96 of 144
KVQA (Knowledge-aware VQA)
It contains manually verified 183K question-answer pairs about more than 18K persons and 24K images.
31 papers · 0 benchmarks
VQA-HAT (Human ATtention) is a dataset to evaluate the informative regions of an image depending on the question being asked about it.
27 papers · 0 benchmarks
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
Social-IQ is an unconstrained benchmark specifically designed to train and evaluate socially intelligent technologies.
23 papers · 0 benchmarks
TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA.
22 papers · 1 benchmark
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
InfiMM-Eval (Complex Open-ended Reasoning Evaluation for Multi-Modal Language Models)
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence.
19 papers · 1 benchmark
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
We collect a new dataset of human-posed free-form natural language questions about CLEVR images.
18 papers · 1 benchmark
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
LIVE-YT-HFR comprises of 480 videos having 6 different frame rates, obtained from 16 diverse contents.
14 papers · 1 benchmark
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
Super-CLEVR is a dataset for Visual Question Answering (VQA) where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently.
13 papers · 0 benchmarks
The VQA-CP dataset was constructed by reorganizing VQA v2 such that the correlation between the question type and correct answer differs in the training and test splits.
13 papers · 1 benchmark
ViQuAE is a dataset for KVQAE (Knowledge-based Visual Question Answering about named Entities), a task which consists in answering questions about named entities grounded in a visual context using a Knowledge Base.
13 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)
Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles.
12 papers · 1 benchmark
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
MP-DocVQA (Multipage Document Visual Question Answering)
The dataset is aimed to perform Visual Question Answering on multipage industry scanned documents.
10 papers · 0 benchmarks
PointQA is a set of datasets for Visual Question Datasets (VQA) that require a pointer to an object in the image to be answered correctly.
10 papers · 0 benchmarks
The question-answer (QA) pairs are automatically generated using state-of-the-art question generation methods based on paintings and comments provided in an existing art understanding dataset.
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LEAF-QA, a comprehensive dataset of 250,000 densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and semantics of these charts.
9 papers · 0 benchmarks
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
SciGraphQA is a large-scale, open-domain dataset focused on generating multi-turn conversational question-answering dialogues centered around understanding and describing scientific graphs and figures.
8 papers · 0 benchmarks
A configurable visual question and answer dataset (COG) to parallel experiments in humans and animals.
7 papers · 0 benchmarks
IQUAD (Interactive Question Answering Dataset)
IQUAD is a dataset for Visual Question Answering in interactive environments.
7 papers · 0 benchmarks
This dataset provides a new split of VQA v2 (similarly to VQA-CP v2), which is built of questions that are hard to answer for biased models.
7 papers · 1 benchmark
DocCVQA (Document Collection Visual Question Answering)
DocCVQA is a Document Visual Question Answering dataset, where the questions are posed over a whole collection of 14,362 scanned documents.
6 papers · 0 benchmarks
Rad-ReStruct is a fine-grained structured reporting dataset for Chest X-Ray images.
6 papers · 0 benchmarks
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
AI2D-RST is a multimodal corpus of 1000 English-language diagrams that represent topics in primary school natural sciences, such as food webs, life cycles, moon phases and human physiology.
5 papers · 0 benchmarks
Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects.
5 papers · 1 benchmark
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
TutorialVQA is a new type of dataset used to find answer spans in tutorial videos.
5 papers · 0 benchmarks
VQA-VS (a new VQA benchmark considering Varying Shortcuts)
The current OOD benchmark VQA-CP v2 only considers one type of shortcut (from question type to answer) and thus still cannot guarantee that the modelrelies on the intended solution rather than a solution specific to this shortcut.
5 papers · 0 benchmarks
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
The VizWiz-VQA-Grounding dataset is a dataset that visually grounds answers to visual questions asked by people with visual impairments.
4 papers · 0 benchmarks
FunQA is a challenging video question answering (QA) dataset specifically designed to evaluate and enhance the depth of video reasoning based on counter-intuitive and fun videos.
3 papers · 0 benchmarks
IllusionVQA is a Visual Question Answering (VQA) dataset with two sub-tasks.
3 papers · 2 benchmarks
MCVQA (Multilingual and Code-mixed Visual Question Answering)
The MCVQA dataset consists of 248, 349 training questions and 121, 512 validation questions for real images in Hindi and Code-mixed.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.