Datasets › NewsQA

NewsQA

Introduced by Adam Trischler et al. in NewsQA: A Machine Comprehension Dataset1 Jan 2017 archive 2025-07-28

The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.

  • Documents are CNN news articles.
  • Questions are written by human users in natural language.
  • Answers may be multiword passages of the source text.
  • Questions may be unanswerable.
  • NewsQA is collected using a 3-stage, siloed process.
  • Questioners see only an article’s headline and highlights.
  • Answerers see the question and the full article, then select an answer passage.
  • Validators see the article, the question, and a set of answers that they rank.
  • NewsQA is more natural and more challenging than previous datasets.

Source: https://www.microsoft.com/en-us/research/project/newsqa-dataset/ Image Source: Trischler et al

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering NewsQA OpenAI/o3-2025-01-31-high EM 92.52 o3-mini vs DeepSeek-R1: Which One is Safer? trust4ai/astral 18 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 272. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
o3-mini vs DeepSeek-R1: Which One is Safer? 1 1 30 Jan 2025 not harvested
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning 4 1 22 Jan 2025 not harvested
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data 0 1 3 Oct 2024 not harvested
Claude 3.5 Sonnet Model Card Addendum 0 1 24 Jun 2024 not harvested
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 1 1 8 Mar 2024 not harvested
DyREx: Dynamic Query Representation for Extractive Question Answering 1 1 26 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
0/1 Deep Neural Networks via Block Coordinate Descent 0 1 19 Jun 2022 not harvested
Time-series Transformer Generative Adversarial Networks 5 1 23 May 2022 not harvested
LinkBERT: Pretraining Language Models with Document Links 1 1 29 Mar 2022 ran 0 of 14 samples (14 unverified)
XAI for Transformers: Better Explanations through Conservative Propagation 1 1 15 Feb 2022 ran 2 of 12 samples (10 unverified)
Learning to Generate Questions by Learning to Recover Answer-containing Sentences 0 1 1 Aug 2021 not harvested
Thinking Like Transformers 5 1 13 Jun 2021 ran 11 of 18 samples (7 unverified; 5 pointer-only for licence)
SpanBERT: Improving Pre-training by Representing and Predicting Spans 6 1 24 Jul 2019 ran 3 of 15 samples (12 unverified; 4 pointer-only for licence)
Densely Connected Attention Propagation for Reading Comprehension 2 1 10 Nov 2018 not harvested
Efficient and Robust Question Answering from Minimal Context over Documents 1 1 21 May 2018 not harvested
A Question-Focused Multi-Factor Attention Network for Question Answering 1 1 25 Jan 2018 not harvested
Making Neural QA as Simple as Possible but not Simpler 3 1 14 Mar 2017 ran 0 of 2 samples (2 unverified)
DeepSense: A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing 1 1 7 Nov 2016 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • NewsQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections