Datasets › RACE

RACE (ReAding Comprehension dataset from Examinations)

Introduced by Guokun Lai et al. in RACE: Large-scale ReAding Comprehension Dataset From Examinations1 Jan 2017 archive 2025-07-28

The ReAding Comprehension dataset from Examinations (RACE) dataset is a machine reading comprehension dataset consisting of 27,933 passages and 97,867 questions from English exams, targeting Chinese students aged 12-18. RACE consists of two subsets, RACE-M and RACE-H, from middle school and high school exams, respectively. RACE-M has 28,293 questions and RACE-H has 69,574. Each question is associated with 4 candidate answers, one of which is correct. The data generation process of RACE differs from most machine reading comprehension datasets - instead of generating questions and answers by heuristics or crowd-sourcing, questions in RACE are specifically designed for testing human reading skills, and are created by domain experts.

Source: Dynamic Fusion Networks for Machine Reading Comprehension Image Source: Lai et al

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 412. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Orca 2: Teaching Small Language Models How to Reason 0 2 18 Nov 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Hierarchical Learning for Generation with Long Source Sequences 0 1 15 Apr 2021 not harvested
Improving Machine Reading Comprehension with Single-choice Decision and Transfer Learning 0 1 6 Nov 2020 not harvested
A BERT-based Distractor Generation Scheme with Multi-tasking and Negative Answer Training Strategies 1 1 12 Oct 2020 not harvested
Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing 3 1 5 Jun 2020 ran 2 of 2 samples (0 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Language Models are Few-Shot Learners 67 4 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
DUMA: Reading Comprehension with Transposition Thinking 3 1 26 Jan 2020 not harvested
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism 10 2 17 Sep 2019 ran 12 of 47 samples (35 unverified; 15 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
XLNet: Generalized Autoregressive Pretraining for Language Understanding 27 2 19 Jun 2019 ran 10 of 24 samples (14 unverified; 3 pointer-only for licence)
Option Comparison Network for Multiple-choice Reading Comprehension 0 1 7 Mar 2019 not harvested
Dual Co-Matching Network for Multi-choice Reading Comprehension 0 1 27 Jan 2019 not harvested
Improving Language Understanding by Generative Pre-Training 13 1 11 Jun 2018 not harvested
Multi-range Reasoning for Machine Comprehension 0 1 24 Mar 2018 not harvested

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • RACE

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections