Datasets › Natural Questions

Natural Questions

Introduced by Tom Kwiatkowski et al. in Natural Questions: a Benchmark for Question Answering Research1 Jan 2019 archive 2025-07-28

The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples. Each example is comprised of a google.com query and a corresponding Wikipedia page. Each Wikipedia page has a passage (or long answer) annotated on the page that answers the question and one or more short spans from the annotated passage containing the actual answer. The long and the short answer annotations can however be empty. If they are both empty, then there is no answer on the page at all. If the long answer annotation is non-empty, but the short answer annotation is empty, then the annotated passage answers the question but no explicit short answer could be found. Finally 1% of the documents have a passage annotated with a short answer that is “yes” or “no”, instead of a list of short spans.

Source: A BERT Baseline for the Natural Questions Image Source: https://paperswithcode.com/paper/natural-questions-a-benchmark-for-question/

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 49 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,404. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Search-o1: Agentic Search-Enhanced Large Reasoning Models 2 1 9 Jan 2025 not harvested
BM25S: Orders of magnitude faster lexical search via eager sparse scoring 3 4 4 Jul 2024 ran 10 of 22 samples (12 unverified)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 0 4 2 Jul 2024 not harvested
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation 1 1 26 Jun 2024 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers 1 2 22 Mar 2024 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG 0 2 18 Jan 2024 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 1 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
REPLUG: Retrieval-Augmented Black-Box Language Models 3 2 30 Jan 2023 ran 0 of 13 samples (13 unverified; 13 pointer-only for licence)
Retrieval as Attention: End-to-end Learning of Retrieval and Reading within a Single Transformer 1 2 5 Dec 2022 not harvested
FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering 0 1 18 Nov 2022 not harvested
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 ran 2 of 2 samples (0 unverified)
Atlas: Few-shot Learning with Retrieval Augmented Language Models 2 4 5 Aug 2022 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
No Parameter Left Behind: How Distillation and Model Size Affect Zero-Shot Retrieval 1 1 6 Jun 2022 not harvested
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Augmenting Document Representations for Dense Retrieval with Interpolation and Perturbation 1 1 15 Mar 2022 not harvested
SGPT: GPT Sentence Embeddings for Semantic Search 1 2 17 Feb 2022 ran 1 of 1 samples (0 unverified)
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 0 3 13 Dec 2021 not harvested
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
Improving language models by retrieving from trillions of tokens 2 1 8 Dec 2021 ran 16 of 23 samples (7 unverified; 3 pointer-only for licence)
Salient Phrase Aware Dense Retrieval: Can a Dense Retriever Imitate a Sparse One? 2 1 13 Oct 2021 not harvested
R2-D2: A Modular Baseline for Open-Domain Question Answering 1 5 8 Sep 2021 not harvested
0.8% Nyquist computational ghost imaging via non-experimental deep learning 0 2 17 Aug 2021 not harvested
Domain-matched Pre-training Tasks for Dense Retrieval 1 1 28 Jul 2021 not harvested
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering 2 2 9 Jun 2021 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Efficient Passage Retrieval with Hashing for Open-domain Question Answering 1 2 2 Jun 2021 not harvested
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models 3 2 17 Apr 2021 ran 3 of 4 samples (1 unverified)

The full list of 49 is in the JSON twin.

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Natural Questions
  • Natural Questions (long)
  • Natural Questions (short)
  • NQ
  • NQ (BEIR)

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections