Datasets › DocVQA

DocVQA

Introduced by Minesh Mathew et al. in DocVQA: A Dataset for VQA on Document Images archive 2025-07-28

DocVQA consists of 50,000 questions defined on 12,000+ document images.

Source: DocVQA: A Dataset for VQA on Document Images

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering (VQA) DocVQA test Human ANLS 0.9436 DocVQA: A Dataset for VQA on Document Images anisha2102/docvqa +2 33 Compare
Visual Question Answering (VQA) DocVQA val BERT LARGE Baseline Accuracy 54.48 DocVQA: A Dataset for VQA on Document Images anisha2102/docvqa +2 2 Compare
Visual Question Answering (VQA) DocVQA ChatGPT 3.5 with LAPDoc Prompt (SpatialFormat) ANLS 79.8 LAPDoc: Layout-Aware Prompting for Documents — 1 Compare

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 290. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Multi-label Cluster Discrimination for Visual Representation Learning 1 1 24 Jul 2024 ran 7 of 11 samples (4 unverified)
LAPDoc: Layout-Aware Prompting for Documents 0 1 15 Feb 2024 not harvested
ScreenAI: A Vision-Language Model for UI and Infographics Understanding 2 1 7 Feb 2024 not harvested
Omni-SMoLA: Boosting Generalist Multimodal Models with Soft Mixture of Low-rank Experts 0 2 1 Dec 2023 not harvested
PaLI-3 Vision Language Models: Smaller, Faster, Stronger 1 2 13 Oct 2023 ran 3 of 3 samples (0 unverified)
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 2 3 24 Aug 2023 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
DocFormerv2: Local Features for Document Understanding 1 1 2 Jun 2023 not harvested
Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering 3 3 1 Jun 2023 ran 1 of 10 samples (9 unverified)
PaLI-X: On Scaling up a Multilingual Vision and Language Model 2 3 29 May 2023 ran 6 of 7 samples (1 unverified)
DUBLIN -- Document Understanding By Language-Image Network 0 2 23 May 2023 not harvested
MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering 1 1 19 Dec 2022 not harvested
Unifying Vision, Text, and Layout for Universal Document Processing 5 2 5 Dec 2022 ran 4 of 17 samples (13 unverified; 3 pointer-only for licence)
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding 2 2 12 Oct 2022 ran 2 of 7 samples (5 unverified)
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding 4 2 7 Oct 2022 ran 0 of 5 samples (5 unverified)
End-to-end Document Recognition and Understanding with Dessurt 2 1 30 Mar 2022 ran 0 of 11 samples (11 unverified)
OCR-free Document Understanding Transformer 5 1 30 Nov 2021 ran 0 of 9 samples (9 unverified)
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer 1 2 18 Feb 2021 not harvested
LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding 9 2 29 Dec 2020 not harvested
DocVQA: A Dataset for VQA on Document Images 3 4 1 Jul 2020 ran 2 of 6 samples (4 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • DocVQA val
  • DocVQA test
  • DocVQA

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections