Datasets › DROP

DROP (Discrete Reasoning Over Paragraphs)

Introduced by Dheeru Dua et al. in DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs1 Jan 2019 archive 2025-07-28

Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over them (such as addition, counting, or sorting). These operations require a much more comprehensive understanding of the content of paragraphs than what was necessary for prior datasets. The questions consist of passages extracted from Wikipedia articles. The dataset is split into a training set of about 77,000 questions, a development set of around 9,500 questions and a hidden test set similar in size to the development set.

Source: https://allennlp.org/drop Image Source: DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering DROP Test QDGAT (ensemble) F1 88.38 Question Directed Graph Attention Network for Numerical... — 16 Compare
Question Answering DROP PaLM 540B (Self Improvement, Self Consistency) Accuracy 83 Large Language Models Can Self-Improve — 6 Compare
Text Generation Drop (3-Shot) no rows — — 0 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 382. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Orca 2: Teaching Small Language Models How to Reason 0 2 18 Nov 2023 not harvested
PaLM 2 Technical Report 1 1 17 May 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
Large Language Models Can Self-Improve 0 6 20 Oct 2022 not harvested
Reasoning Like Program Executors 1 1 27 Jan 2022 not harvested
Question Directed Graph Attention Network for Numerical Reasoning over Text 0 1 16 Sep 2020 not harvested
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Neural Symbolic Reader: Scalable Integration of Distributed and Symbolic Representations for Reading Comprehension 0 1 1 May 2020 not harvested
Injecting Numerical Reasoning Skills into Language Models 2 1 9 Apr 2020 not harvested
NumNet: Machine Reading Comprehension with Numerical Reasoning 2 1 15 Oct 2019 not harvested
A Simple and Effective Model for Answering Multi-span Questions 4 1 29 Sep 2019 not harvested
Giving BERT a Calculator: Finding Operations and Arguments with Reading Comprehension 0 1 31 Aug 2019 not harvested
A Multi-Type Multi-Span Network for Reading Comprehension that Requires Discrete Reasoning 1 1 15 Aug 2019 not harvested
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs 3 2 1 Mar 2019 not harvested

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • DROP Test
  • DROP
  • Drop (3-Shot)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections