Home › Datasets › task › Reading Comprehension

Reading Comprehension datasets

archive 2025-07-28

95 datasets carry the task tag "Reading Comprehension" (the task itself: Reading Comprehension), ordered by the archive's paper count. Page 1 of 2: 48 shown of 95. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Reading Comprehension datasets 1–48 of 95

MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
RACE (ReAding Comprehension dataset from Examinations)
The ReAding Comprehension dataset from Examinations (RACE) dataset is a machine reading comprehension dataset consisting of 27,933 passages and 97,867 questions from English exams, targeting Chinese students aged 12-18.
412 papers · 3 benchmarks
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
CoQA (Conversational Question Answering Challenge)
CoQA is a large-scale dataset for building Conversational Question Answering systems.
281 papers · 2 benchmarks
The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.
272 papers · 1 benchmark
The NarrativeQA dataset includes a list of documents with Wikipedia summaries, links to full stories, and questions and answers.
206 papers · 1 benchmark
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others.
183 papers · 1 benchmark
QuAC (Question Answering in Context)
Question Answering in Context is a large-scale dataset that consists of around 14K crowdsourced Question Answering dialogs with 98K question-answer pairs in total.
178 papers · 1 benchmark
MultiRC (Multi-Sentence Reading Comprehension)
MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph.
162 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
TyDiQA (Typologically Diverse Question Answering)
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 200K question-answer pairs.
148 papers · 0 benchmarks
A new open-vocabulary language modelling benchmark derived from books.
147 papers · 1 benchmark
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
QASC (Question Answering via Sentence Composition)
QASC is a question-answering dataset with a focus on sentence composition.
114 papers · 0 benchmarks
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
CLUE (Chinese Language Understanding Evaluation Benchmark)
CLUE is a Chinese Language Understanding Evaluation benchmark.
99 papers · 8 benchmarks
CBT (Children’s Book Test)
Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context.
92 papers · 1 benchmark
Logical reasoning is an important ability to examine, analyze, and critically evaluate arguments as they occur in ordinary language as the definition from Law School Admission Council.
88 papers · 4 benchmarks
OPUS-100 is an English-centric multilingual corpus covering 100 languages.
78 papers · 0 benchmarks
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
DREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset.
68 papers · 2 benchmarks
DuReader is a large-scale open-domain Chinese machine reading comprehension dataset.
65 papers · 1 benchmark
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
QReCC contains 14K conversations with 81K question-answer pairs.
51 papers · 0 benchmarks
NomBank is an annotation project at New York University that is related to the PropBank project at the University of Colorado.
50 papers · 0 benchmarks
Quoref is a QA dataset which tests the coreferential reasoning capability of reading comprehension systems.
50 papers · 0 benchmarks
emrQA has 1 million question-logical form and 400,000+ questionanswer evidence pairs.
50 papers · 0 benchmarks
CMRC 2018 (Chinese Machine Reading Comprehension 2018)
CMRC 2018 is a dataset for Chinese Machine Reading Comprehension.
49 papers · 0 benchmarks
TQA (Textbook Question Answering)
The TextbookQuestionAnswering (TQA) dataset is drawn from middle school science curricula.
48 papers · 1 benchmark
QUASAR (QUestion Answering by Search And Reading)
The Question Answering by Search And Reading (QUASAR) is a large-scale dataset consisting of QUASAR-S and QUASAR-T.
47 papers · 1 benchmark
CoS-E (Commonsense Explanations Dataset)
CoS-E consists of human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations Source: Explain Yourself!
46 papers · 0 benchmarks
ShARC (Shaping Answers with Rules through Conversation)
ShARC is a Conversational Question Answering dataset focussing on question answering from texts containing rules.
43 papers · 0 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
The ProPara dataset is designed to train and test comprehension of simple paragraphs describing processes (e.g., photosynthesis), designed for the task of predicting, tracking, and answering questions about how entities change during the…
34 papers · 0 benchmarks
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
Holl-E is a dataset containing movie chats wherein each response is explicitly generated by copying and/or modifying sentences from unstructured background knowledge such as plots, comments and reviews about the movie.
30 papers · 0 benchmarks
A machine reading comprehension (MRC) dataset with discourse structure built over multiparty dialog.
29 papers · 2 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks
IIRC (Incomplete Information Reading Comprehension)
Contains more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.
26 papers · 0 benchmarks
WikiReading is a large-scale natural language understanding task and publicly-available dataset with 18 million instances.
26 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.