Home › Datasets › task › Visual Question Answering

Visual Question Answering datasets

archive 2025-07-28

32 datasets carry the task tag "Visual Question Answering" (the task itself: Visual Question Answering), ordered by the archive's paper count. Page 1 of 1: 32 shown of 32. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Visual Question Answering datasets 1–32 of 32

The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs.
749 papers · 8 benchmarks
CLEVR (Compositional Language and Elementary Visual Reasoning)
CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset.
657 papers · 3 benchmarks
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
MMBench is a multi-modality benchmark.
384 papers · 1 benchmark
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images.
366 papers · 7 benchmarks
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
339 papers · 2 benchmarks
VizWiz (VizWiz-VQA)
The VizWiz-VQA dataset originates from a natural visual question answering setting where blind people each took an image and recorded a spoken question about it, together with 10 crowdsourced answers per visual question.
260 papers · 7 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset.
66 papers · 5 benchmarks
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
17 papers · 1 benchmark
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)
Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles.
12 papers · 1 benchmark
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
SciGraphQA is a large-scale, open-domain dataset focused on generating multi-turn conversational question-answering dialogues centered around understanding and describing scientific graphs and figures.
8 papers · 0 benchmarks
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
EarthVQA (A multi-modal multi-task VQA dataset for remote sensing)
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning.
6 papers · 1 benchmark
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
EVJVQA (English-Japanese-Vietnamese Visual Question Answering)
EVJVQA, the first multilingual Visual Question Answering dataset with three languages: English, Vietnamese, and Japanese, is released in this task.
4 papers · 0 benchmarks
OpenViVQA (Open-domain Visual Question Answering in Vietnamese)
In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for blind people, or…
4 papers · 0 benchmarks
Kvasir-VQA (A Text-Image Pair GI Tract Dataset)
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
The dataset was created to address the crucial need for effective Extreme Weather Events Detection (EWED), an increasingly urgent task due to the rising frequency of such events driven by global warming.
2 papers · 0 benchmarks
GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.
2 papers · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
MapEval-Visual contains 400 image-question-answer triplets.
1 paper · 1 benchmark
ViLCo (ViLCo-Bench)
We propose the first standardized benchmark in multimodal continual learning for video data, defining protocols for training and metrics for evaluation.
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
uBench (MicroBench)
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.