Browse State-of-the-Art › Visual Question Answering (VQA)

Visual Question Answering (VQA)

1,039 papers with code · 76 benchmarks · 144 datasets archive 2025-07-28

Computer VisionNatural Language Processing

Visual Question Answering (VQA) is a task in computer vision that involves answering questions about an image. The goal of VQA is to teach machines to understand the content of an image and answer questions about it in natural language.

Image Source: visualqa.org

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

76 leaderboard tables shown for this task, 76 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 76 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
GQA Test2019 (127 rows) human — — — Compare
VQA v2 test-dev (56 rows) PaLI PaLI: A Jointly-Scaled Multilingual Language-Image Model code Syntology ran 2 of 4 samples · 2 unverified Compare
VQA v2 test-std (38 rows) BEiT-3 Image as a Foreign Language: BEiT Pretraining for All Vision and... code — Compare
OK-VQA (37 rows) PaLI-X-VPD Visual Program Distillation: Distilling Tools and Programmatic... — — Compare
MSVD-QA (36 rows) VLAB VLAB: Enhancing Video Language Pre-training by Feature Adapting... — — Compare
MSRVTT-QA (34 rows) VLAB VLAB: Enhancing Video Language Pre-training by Feature Adapting... — — Compare
DocVQA test (33 rows) Human DocVQA: A Dataset for VQA on Document Images code Syntology ran 2 of 6 samples · 4 unverified Compare
InfographicVQA (21 rows) Gemini Ultra (pixel only) Gemini: A Family of Highly Capable Multimodal Models code — Compare
GQA test-dev (17 rows) CFR Coarse-to-Fine Reasoning for Visual Question Answering code Syntology ran 4 of 9 samples · 5 unverified Compare
VizWiz 2020 VQA (16 rows) PaLI PaLI: A Jointly-Scaled Multilingual Language-Image Model code Syntology ran 2 of 4 samples · 2 unverified Compare
A-OKVQA (15 rows) SMoLA-PaLI-X Specialist Model Omni-SMoLA: Boosting Generalist Multimodal Models with Soft... — — Compare
CLEVR (15 rows) NS-VQA (1K programs) Neural-Symbolic VQA: Disentangling Reasoning from Vision and... code Syntology ran 2 of 2 samples · 0 unverified Compare
COCO Visual Question Answering (VQA) real images 1.0 open ended (14 rows) MCB 7 att. Multimodal Compact Bilinear Pooling for Visual Question Answering... code — Compare
InfiMM-Eval (14 rows) GPT-4V GPT-4 Technical Report code Syntology ran 2 of 5 samples · 3 unverified Compare
IconQA (12 rows) Patch-TRM IconQA: A New Benchmark for Abstract Diagram Understanding and... code Syntology ran 8 of 10 samples · 2 unverified Compare
TextVQA test-standard (12 rows) PaLI PaLI: A Jointly-Scaled Multilingual Language-Image Model code Syntology ran 2 of 4 samples · 2 unverified Compare
VCR (Q-A) test (11 rows) GPT4RoI GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest code Syntology ran 3 of 5 samples · 2 unverified Compare
VQA v2 val (11 rows) BLIP-2 ViT-G FlanT5 XXL (zero-shot) BLIP-2: Bootstrapping Language-Image Pre-training with Frozen... code Syntology ran 4 of 8 samples · 4 unverified Compare
COCO Visual Question Answering (VQA) real images 1.0 multiple choice (10 rows) MCB 7 att. Multimodal Compact Bilinear Pooling for Visual Question Answering... code — Compare
VizWiz 2018 (10 rows) LXR955, No Ensemble LXMERT: Learning Cross-Modality Encoder Representations from Transformers code Syntology ran 4 of 15 samples · 11 unverified Compare
VQA-CP (10 rows) CSS Counterfactual Samples Synthesizing for Robust Visual Question Answering code Syntology ran 3 of 3 samples · 0 unverified Compare
VQA-CE (9 rows) RandImg Beyond Question-Based Biases: Assessing Multimodal Shortcut... code Syntology ran 1 of 2 samples · 1 unverified Compare
VLM2-Bench (9 rows) GPT-4o GPT-4o System Card — — Compare
VCR (QA-R) test (8 rows) GPT4RoI GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest code Syntology ran 3 of 5 samples · 2 unverified Compare
GQA test-std (7 rows) ProTo ProTo: Program-Guided Transformer for Program-Guided Tasks code Syntology ran 0 of 1 samples · 1 unverified Compare
VCR (Q-AR) test (7 rows) GPT4RoI GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest code Syntology ran 3 of 5 samples · 2 unverified Compare
VQA v1 test-dev (7 rows) SAAA (ResNet) Show, Ask, Attend, and Answer: A Strong Baseline For Visual... code Syntology ran 9 of 9 samples · 0 unverified Compare
IllusionVQA (7 rows) GPT4-Vision 4-shot IllusionVQA: A Challenging Optical Illusion Dataset for Vision... code Syntology ran 5 of 5 samples · 0 unverified Compare
InfoSeek (7 rows) RA-VQAv2 w/ PreFLMR PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers code — Compare
VizWiz 2020 Answerability (6 rows) CLIP-Ensemble Less Is More: Linear Layers on CLIP Features as Powerful VizWiz Model — — Compare
VQA v1 test-std (6 rows) SAAA (ResNet) Show, Ask, Attend, and Answer: A Strong Baseline For Visual... code Syntology ran 9 of 9 samples · 0 unverified Compare
WHOOPS! (6 rows) BLIP2 FlanT5-XXL (Fine-tuned) Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of... — — Compare
CLEVR-Humans (5 rows) MDETR MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding code Syntology ran 6 of 11 samples · 5 unverified Compare
QLEVR (5 rows) MAC QLEVR: A Diagnostic Dataset for Quantificational Language and... code — Compare
AutoHallusion (5 rows) GPT-4V AutoHallusion: Automatic Generation of Hallucination Benchmarks... code Syntology ran 1 of 1 samples · 0 unverified Compare
COCO Visual Question Answering (VQA) real images 2.0 open ended (4 rows) HDU-USYD-UNCC VQA: Visual Question Answering code Syntology ran 6 of 7 samples · 1 unverified Compare
COCO Visual Question Answering (VQA) abstract images 1.0 open ended (4 rows) Graph VQA Graph-Structured Representations for Visual Question Answering — — Compare
COCO Visual Question Answering (VQA) abstract 1.0 multiple choice (4 rows) Graph VQA Graph-Structured Representations for Visual Question Answering — — Compare
PlotQA-D1 (4 rows) MatCha4096 + LaMenDa Synthesize Step-by-Step: Tools Templates and LLMs as Data... — — Compare
PlotQA-D2 (4 rows) MatCha4096 + LaMenDa Synthesize Step-by-Step: Tools Templates and LLMs as Data... — — Compare
Visual7W (4 rows) CMN Modeling Relationships in Referential Expressions with... code — Compare
HallusionBench (4 rows) GPT-4V HallusionBench: An Advanced Diagnostic Suite for Entangled... code Syntology ran 2 of 8 samples · 6 unverified Compare
AI2D (4 rows) SMoLA-PaLI-X Specialist Model Omni-SMoLA: Boosting Generalist Multimodal Models with Soft... — — Compare
PMC-VQA (4 rows) MedVInT PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering code Syntology ran 3 of 5 samples · 2 unverified Compare
F-VQA (3 rows) ZS-F-VQA Zero-shot Visual Question Answering using Knowledge Graph code — Compare
FigureQA - test 1 (3 rows) PReFIL Answering Questions about Data Visualizations using Efficient... code — Compare
VCR (Q-A) dev (3 rows) VL-BERTLARGE VL-BERT: Pre-training of Generic Visual-Linguistic Representations code Syntology ran 0 of 1 samples · 1 unverified Compare
VCR (Q-AR) dev (3 rows) VL-BERTLARGE VL-BERT: Pre-training of Generic Visual-Linguistic Representations code Syntology ran 0 of 1 samples · 1 unverified Compare
VCR (QA-R) dev (3 rows) VL-BERTLARGE VL-BERT: Pre-training of Generic Visual-Linguistic Representations code Syntology ran 0 of 1 samples · 1 unverified Compare
DocVQA val (2 rows) BERT LARGE Baseline DocVQA: A Dataset for VQA on Document Images code Syntology ran 2 of 6 samples · 4 unverified Compare
GQA (2 rows) PEVL+ PEVL: Position-enhanced Pre-training and Prompt Tuning for... code Syntology ran 3 of 7 samples · 4 unverified Compare
GRIT (2 rows) Unified-IOXL Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks — — Compare
TDIUC (2 rows) Accuracy MUREL: Multimodal Relational Reasoning for Visual Question Answering code — Compare
TGIF-QA (2 rows) HiTeA HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training — — Compare
VQA-X (2 rows) OFA-X-MT Harnessing the Power of Multi-Task Pretraining for Ground-Truth... code Syntology ran 1 of 8 samples · 7 unverified Compare
Visual Genome (subjects) (1 row) CMN Modeling Relationships in Referential Expressions with... code — Compare
Visual Genome (pairs) (1 row) CMN Modeling Relationships in Referential Expressions with... code — Compare
VizWiz 2018 Answerability (1 row) ensemble_two_best — — — Compare
ZS-F-VQA (1 row) SAN † - hard mask Zero-shot Visual Question Answering using Knowledge Graph code — Compare
ActivityNet (1 row) BLIP-2 T5 Open-ended VQA benchmarking of Vision-Language models by... code Syntology ran 2 of 2 samples · 0 unverified Compare
ArtQuest (1 row) PrefixLM with CLIP and T5 ArtQuest: Countering Hidden Language Biases in ArtVQA code — Compare
COCO (1 row) InstructBLIP Vicuna Open-ended VQA benchmarking of Vision-Language models by... code Syntology ran 2 of 2 samples · 0 unverified Compare
CORE-MM (1 row) GPT-4V GPT-4 Technical Report code Syntology ran 2 of 5 samples · 3 unverified Compare
DeepForm (1 row) DUBLIN DUBLIN -- Document Understanding By Language-Image Network — — Compare
DocVQA (1 row) ChatGPT 3.5 with LAPDoc Prompt (SpatialFormat) LAPDoc: Layout-Aware Prompting for Documents — — Compare
DVQA test-familiar (1 row) PReFIL (Oracle OCR) Answering Questions about Data Visualizations using Efficient... code — Compare
EgoSchema (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
ImageNet (1 row) BLIP-2 OPT Open-ended VQA benchmarking of Vision-Language models by... code Syntology ran 2 of 2 samples · 0 unverified Compare
MM-Vet (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
MME (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
MVBench (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
OVAD benchmark (1 row) BLIP Open-ended VQA benchmarking of Vision-Language models by... code Syntology ran 2 of 2 samples · 0 unverified Compare
RetVQA (1 row) MI-BART Answer Mining from a Pool of Images: Towards Retrieval-Based... code Syntology ran 6 of 7 samples · 1 unverified Compare
TextVQA (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
Video MME (1 row) Lyra-Pro Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition code Syntology ran 3 of 19 samples · 16 unverified Compare
WebSRC (1 row) DUBLIN DUBLIN -- Document Understanding By Language-Image Network — — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

144 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 144 until expanded.

Subtasks archive 2025-07-28

9 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 1,039 papers with code (2,167 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections