Datasets › SQuAD

SQuAD (Stanford Question Answering Dataset)

Introduced by François Bienvenu et al. in 0-1 laws for pattern occurrences in phylogenetic trees and networks7 Feb 2024 archive 2025-07-28

The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles. In SQuAD, the correct answers of questions can be any sequence of tokens in the given text. Because the questions and answers are produced by humans through crowdsourcing, it is more diverse than some other question-answering datasets. SQuAD 1.1 contains 107,785 question-answer pairs on 536 articles. SQuAD2.0 (open-domain SQuAD, SQuAD-Open), the latest version, combines the 100,000 questions in SQuAD1.1 with over 50,000 un-answerable questions written adversarially by crowdworkers in forms that are similar to the answerable ones.

Source: Deep Learning Based Text Classification: A Comprehensive Review Image Source: https://rajpurkar.github.io/SQuAD-explorer/explore/v2.0/dev/Prime_number.html

Benchmarks archive 2025-07-28

All 12 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 82 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 151. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
GOLD: Generalized Knowledge Distillation via Out-of-Distribution-Guided Language Data Generation 0 1 28 Mar 2024 not harvested
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers 1 2 22 Mar 2024 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
Prompt2Model: Generating Deployable Models from Natural Language Instructions 1 1 23 Aug 2023 not harvested
TextBox 2.0: A Text Generation Library with Pre-trained Language Models 1 2 26 Dec 2022 ran 1 of 1 samples (0 unverified)
DyREx: Dynamic Query Representation for Extractive Question Answering 1 1 26 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback 2 1 22 Oct 2022 not harvested
Enhancing Pre-trained Models with Text Structure Knowledge for Question Generation 0 1 9 Sep 2022 not harvested
LinkBERT: Pretraining Language Models with Document Links 1 1 29 Mar 2022 ran 0 of 14 samples (14 unverified)
ZeroGen: Efficient Zero-shot Learning via Dataset Generation 3 1 16 Feb 2022 ran 3 of 7 samples (4 unverified; 7 pointer-only for licence)
Information Theoretic Representation Distillation 1 2 1 Dec 2021 not harvested
Prune Once for All: Sparse Pre-Trained Language Models 2 9 10 Nov 2021 not harvested
Ensemble ALBERT on SQuAD 2.0 1 1 19 Oct 2021 not harvested
Fine-tune the Entire RAG Architecture (including DPR retriever) for Question-Answering 2 1 22 Jun 2021 not harvested
Pay Attention to MLPs 20 1 17 May 2021 ran 34 of 44 samples (10 unverified; 11 pointer-only for licence)
A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes 0 1 12 Feb 2021 not harvested
EfficientQA : a RoBERTa Based Phrase-Indexed Question-Answering System 0 1 6 Jan 2021 not harvested
Learning to Generate Questions by Recovering Answer-containing Sentences 0 2 1 Jan 2021 not harvested
Learning Dense Representations of Phrases at Scale 4 1 23 Dec 2020 not harvested
LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention 9 7 2 Oct 2020 ran 3 of 10 samples (7 unverified)
SPARTA: Efficient Open-Domain Question Answering via Sparse Transformer Matching Retrieval 1 1 28 Sep 2020 not harvested
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs 1 2 28 May 2020 ran 0 of 16 samples (16 unverified)
Harvesting and Refining Question-Answer Pairs for Unsupervised QA 1 2 6 May 2020 not harvested
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training 3 1 28 Feb 2020 ran 1 of 1 samples (0 unverified)
Retrospective Reader for Machine Reading Comprehension 2 4 27 Jan 2020 not harvested
ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation 5 1 26 Jan 2020 not harvested
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training 5 1 13 Jan 2020 not harvested
Dice Loss for Data-imbalanced NLP Tasks 4 2 7 Nov 2019 ran 0 of 3 samples (3 unverified)
A Recurrent BERT-based Model for Question Generation 1 1 1 Nov 2019 not harvested
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension 47 1 29 Oct 2019 ran 22 of 53 samples (31 unverified; 7 pointer-only for licence)

The full list of 82 is in the JSON twin.

Dataset loaders archive 2025-07-28

38 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • squad_bn
  • squad_adversarial
  • qg_squad
  • The Stanford Question Answering Dataset
  • squad_v2
  • SQuAD
  • SQuAD2.0 dev
  • SQuAD2.0
  • SQuAD1.1 dev
  • SQuAD1.1

10 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections