Datasets › MultiRC

MultiRC (Multi-Sentence Reading Comprehension)

Introduced by Daniel Khashabi et al. in Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences1 Jan 2018 archive 2025-07-28

MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph. The dataset was designed with three key challenges in mind: * The number of correct answer-options for each question is not pre-specified. This removes the over-reliance on answer-options and forces them to decide on the correctness of each candidate answer independently of others. In other words, the task is not to simply identify the best answer-option, but to evaluate the correctness of each answer-option individually. * The correct answer(s) is not required to be a span in the text. * The paragraphs in the dataset have diverse provenance by being extracted from 7 different domains such as news, fiction, historical text etc., and hence are expected to be more diverse in their contents as compared to single-domain datasets. The entire corpus consists of around 10K questions (including about 6K multiple-sentence questions). The 60% of the data is released as training and development data. The rest of the data is saved for evaluation and every few months a new unseen additional data is included for evaluation to prevent unintentional overfitting over time.

Source: https://cogcomp.seas.upenn.edu/multirc/ Image Source: https://paperswithcode.com/paper/looking-beyond-the-surface-a-challenge-set/

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering MultiRC PaLM 540B (finetuned) F1 90.1 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 30 Compare

Papers archive 2025-07-28

15 shown of 15 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 162. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
Hungry Hungry Hippos: Towards Language Modeling with State Space Models 3 4 28 Dec 2022 ran 7 of 15 samples (8 unverified)
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE 0 2 4 Dec 2022 not harvested
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 ran 2 of 2 samples (0 unverified)
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model 1 1 2 Aug 2022 ran 1 of 1 samples (0 unverified)
N-Grammer: Augmenting Transformers with latent n-grams 2 1 13 Jul 2022 ran 0 of 6 samples (6 unverified)
PaLM: Scaling Language Modeling with Pathways 7 1 5 Apr 2022 ran 30 of 37 samples (7 unverified)
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 2 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
KELM: Knowledge Enhanced Pre-Trained Language Representations with Message Passing on Hierarchical Relational Graphs 1 1 9 Sep 2021 ran 1 of 6 samples (5 unverified)
Finetuned Language Models Are Zero-Shot Learners 8 3 3 Sep 2021 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 2 23 Oct 2019 ran 2 of 31 samples (29 unverified)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 1 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MultiRC

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections