Datasets › A-OKVQA

A-OKVQA

Introduced by Dustin Schwenk et al. in A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge3 Jun 2022 archive 2025-07-28

A-OKVQA is crowdsourced visual question answering dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering (VQA) A-OKVQA SMoLA-PaLI-X Specialist Model MC Accuracy 83.75 Omni-SMoLA: Boosting Generalist Multimodal Models with... — 15 Compare

Papers archive 2025-07-28

13 shown of 13 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 154. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning 1 1 19 Mar 2024 ran 5 of 8 samples (3 unverified)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models 0 1 5 Dec 2023 not harvested
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts 0 1 1 Dec 2023 not harvested
Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training 1 1 23 Nov 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
A Simple Baseline for Knowledge-Based Visual Question Answering 0 1 20 Oct 2023 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering 1 1 3 Mar 2023 ran 0 of 8 samples (8 unverified)
PromptCap: Prompt-Guided Task-Aware Image Captioning 1 1 15 Nov 2022 not harvested
VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge 1 1 24 Oct 2022 ran 0 of 2 samples (2 unverified)
Webly Supervised Concept Expansion for General Purpose Vision Models 0 1 4 Feb 2022 not harvested
KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA 0 1 20 Dec 2020 not harvested
LXMERT: Learning Cross-Modality Encoder Representations from Transformers 9 1 20 Aug 2019 ran 4 of 15 samples (11 unverified; 3 pointer-only for licence)
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks 11 3 6 Aug 2019 ran 10 of 34 samples (24 unverified; 34 pointer-only for licence)
Pythia v0.1: the Winning Entry to the VQA Challenge 2018 9 1 26 Jul 2018 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • A-OKVQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections