Datasets › Visual Question Answering v2.0

Visual Question Answering v2.0 (VQA v2.0)

Introduced by Yash Goyal et al. in Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering2 Dec 2016 archive 2025-07-28

Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images. These questions require an understanding of vision, language and commonsense knowledge to answer. It is the second version of the VQA dataset.

  • 265,016 images (COCO and abstract scenes)
  • At least 3 questions (5.4 questions on average) per image
  • 10 ground truth answers per question
  • 3 plausible (but likely incorrect) answers per question
  • Automatic evaluation metric

The first version of the dataset was released in October 2015.

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 66 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 366. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts 1 1 9 May 2024 ran 10 of 12 samples (2 unverified)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 2 11 Apr 2024 not harvested
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 2 1 21 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects 0 1 8 Dec 2023 not harvested
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback 4 1 1 Dec 2023 not harvested
LXMERT Model Compression for Visual Question Answering 2 2 23 Oct 2023 not harvested
Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 21 Sep 2023 not harvested
Emu: Generative Pretraining in Multimodality 2 1 11 Jul 2023 ran 0 of 2 samples (2 unverified)
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities 2 2 18 May 2023 ran 2 of 7 samples (5 unverified)
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset 1 2 17 Apr 2023 not harvested
Prismer: A Vision-Language Model with Multi-Task Experts 2 2 4 Mar 2023 not harvested
Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 11 Feb 2023 not harvested
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 4 1 1 Feb 2023 ran 9 of 19 samples (10 unverified)
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 18 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks 1 1 12 Jan 2023 not harvested
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 4 22 Nov 2022 ran 2 of 6 samples (4 unverified; 6 pointer-only for licence)
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training 3 2 17 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 1 14 Sep 2022 ran 2 of 4 samples (2 unverified)
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks 2 2 22 Aug 2022 not harvested
Prompt Tuning for Generative Multimodal Pretrained Models 1 1 4 Aug 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
LaKo: Knowledge-driven Visual Question Answering via Late Knowledge-to-Text Injection 1 1 26 Jul 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Language Models are General-Purpose Interfaces 1 1 13 Jun 2022 not harvested
mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections 3 2 24 May 2022 not harvested
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 ran 9 of 17 samples (8 unverified)
Flamingo: a Visual Language Model for Few-Shot Learning 5 3 29 Apr 2022 ran 18 of 24 samples (6 unverified; 7 pointer-only for licence)
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework 4 2 7 Feb 2022 ran 1 of 1 samples (0 unverified)
Florence: A New Foundation Model for Computer Vision 2 2 22 Nov 2021 not harvested
Achieving Human Parity on Visual Question Answering 0 1 17 Nov 2021 not harvested
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts 1 1 16 Nov 2021 ran 1 of 1 samples (0 unverified)
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation 0 2 16 Nov 2021 not harvested

The full list of 66 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • VQA 2.0
  • VQA v2
  • VQA v2 test-dev
  • VQA v2 test-std
  • Visual Question Answering v2.0
  • VQA v2 val

6 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections