Datasets › OK-VQA

OK-VQA (Outside Knowledge Visual Question Answering)

Introduced by Kenneth Marino et al. in OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge archive 2025-07-28

Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.

Source: OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge Image Source: https://okvqa.allenai.org/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 368. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning 1 1 19 Mar 2024 ran 5 of 8 samples (3 unverified)
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects 0 1 8 Dec 2023 not harvested
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models 0 1 5 Dec 2023 not harvested
A Simple Baseline for Knowledge-Based Visual Question Answering 0 1 20 Oct 2023 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering 1 3 29 Sep 2023 ran 8 of 8 samples (0 unverified; 8 pointer-only for licence)
Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 21 Sep 2023 not harvested
PaLI-X: On Scaling up a Multilingual Vision and Language Model 2 1 29 May 2023 ran 6 of 7 samples (1 unverified)
PaLM-E: An Embodied Multimodal Language Model 2 1 6 Mar 2023 not harvested
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering 1 1 3 Mar 2023 ran 0 of 8 samples (8 unverified)
Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 11 Feb 2023 not harvested
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 6 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory 1 1 10 Dec 2022 not harvested
PromptCap: Prompt-Guided Task-Aware Image Captioning 1 1 15 Nov 2022 not harvested
VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge 1 1 24 Oct 2022 ran 0 of 2 samples (2 unverified)
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training 3 1 17 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Retrieval Augmented Visual Question Answering with Outside Knowledge 1 3 7 Oct 2022 not harvested
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 1 14 Sep 2022 ran 2 of 4 samples (2 unverified)
LaKo: Knowledge-driven Visual Question Answering via Late Knowledge-to-Text Injection 1 2 26 Jul 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Language Models are General-Purpose Interfaces 1 1 13 Jun 2022 not harvested
REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question Answering 1 2 2 Jun 2022 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Flamingo: a Visual Language Model for Few-Shot Learning 5 3 29 Apr 2022 ran 18 of 24 samples (6 unverified; 7 pointer-only for licence)
Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question Answering 0 1 1 Jan 2022 not harvested
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation 0 1 16 Nov 2021 not harvested
A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models 1 1 16 Oct 2021 not harvested
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA 1 1 10 Sep 2021 ran 2 of 2 samples (0 unverified)
Multimodal Few-Shot Learning with Frozen Language Models 0 1 25 Jun 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • OK-VQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections