Home › Datasets › task › Machine Reading Comprehension
Machine Reading Comprehension datasets
archive 2025-07-28
43 datasets carry the task tag "Machine Reading Comprehension" (the task itself: Machine Reading Comprehension), ordered by the archive's paper count. Page 1 of 1: 43 shown of 43. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Machine Reading Comprehension datasets 1–43 of 43
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
CBT (Children’s Book Test)
Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context.
92 papers · 1 benchmark
Logical reasoning is an important ability to examine, analyze, and critically evaluate arguments as they occur in ordinary language as the definition from Law School Admission Council.
88 papers · 4 benchmarks
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
DREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset.
68 papers · 2 benchmarks
MuTual is a retrieval-based dataset for multi-turn dialogue reasoning, which is modified from Chinese high school English listening comprehension test data.
55 papers · 0 benchmarks
C3 is a free-form multiple-Choice Chinese machine reading Comprehension dataset.
54 papers · 0 benchmarks
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
Quoref is a QA dataset which tests the coreferential reasoning capability of reading comprehension systems.
50 papers · 0 benchmarks
CMRC 2018 (Chinese Machine Reading Comprehension 2018)
CMRC 2018 is a dataset for Chinese Machine Reading Comprehension.
49 papers · 0 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
ChID (Chinese IDiom dataset)
ChID is a large-scale Chinese IDiom dataset for cloze test.
38 papers · 0 benchmarks
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
Holl-E is a dataset containing movie chats wherein each response is explicitly generated by copying and/or modifying sentences from unstructured background knowledge such as plots, comments and reviews about the movie.
30 papers · 0 benchmarks
A machine reading comprehension (MRC) dataset with discourse structure built over multiparty dialog.
29 papers · 2 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
With social media becoming increasingly popular on which lots of news and real-time events are reported, developing automated question answering systems is critical to the effectiveness of many applications that rely on real-time knowledge.
19 papers · 1 benchmark
A new dataset for the low-resource language as Vietnamese to evaluate MRC models.
15 papers · 0 benchmarks
A large-scale cloze-style biomedical MRC dataset.
13 papers · 1 benchmark
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
CMRC 2017 (Chinese Machine Reading Comprehension 2017)
Contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set.
11 papers · 0 benchmarks
A human-curated ChineseReading Comprehension dataset on Opinion.
9 papers · 0 benchmarks
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance.
5 papers · 0 benchmarks
RuCoS (Russian Reading Comprehension with Commonsense Reasoning)
Russian reading comprehension with Commonsense reasoning (RuCoS) is a large-scale reading comprehension dataset that requires commonsense reasoning.
5 papers · 1 benchmark
UIT-ViNewsQA is a new corpus for the Vietnamese language to evaluate healthcare reading comprehension models.
5 papers · 0 benchmarks
ViMMRC (Vietnamese Multiple-choice Machine Reading Comprehension Corpus)
A challenging machine comprehension corpus with multiple-choice questions, intended for research on the machine comprehension of Vietnamese text.
4 papers · 0 benchmarks
The UIT-ViWikiQA is a dataset for evaluating sentence extraction-based machine reading comprehension in the Vietnamese language.
3 papers · 0 benchmarks
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
A dataset of around 2 million examples for machine reading-comprehension.
2 papers · 0 benchmarks
MAUD is an expert-annotated merger agreement reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points study, where lawyers and law students answered 92 questions about 152 merger agreements.
2 papers · 0 benchmarks
UQuAD (Urdu Question Answering Dataset)
Large scale machine reading comprehension dataset in Urdu language.
2 papers · 1 benchmark
A newly developed public dataset and the task of multiple property extraction.
2 papers · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
PersianQA (Persian Question Answering Dataset)
PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia.
1 paper · 0 benchmarks
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.