Home › Datasets › task › Reading Comprehension

Reading Comprehension datasets

archive 2025-07-28

95 datasets carry the task tag "Reading Comprehension" (the task itself: Reading Comprehension), ordered by the archive's paper count. Page 2 of 2: 47 shown of 95. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Reading Comprehension datasets 49–95 of 95

MCScript is used as the official dataset of SemEval2018 Task11.
24 papers · 0 benchmarks
ROPES (Reasoning Over Paragraph Effects in Situations)
ROPES is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation.
24 papers · 0 benchmarks
In SpokenSQuAD, the document is in spoken form, the input question is in the form of text and the answer to each question is always a span in the document.
24 papers · 1 benchmark
BookTest is a new dataset similar to the popular Children’s Book Test (CBT), however more than 60 times larger.
22 papers · 0 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
A publicly available dataset with 242k labeled sections in English and German from two distinct domains: diseases and cities.
17 papers · 0 benchmarks
A new dataset for the low-resource language as Vietnamese to evaluate MRC models.
15 papers · 0 benchmarks
With the same format as WikiHop, the MedHop dataset is based on research paper abstracts from PubMed, and the queries are about interactions between pairs of drugs.
14 papers · 0 benchmarks
A large-scale cloze-style biomedical MRC dataset.
13 papers · 1 benchmark
A dataset of ~19K questions that are elicited while a person is reading through a document.
13 papers · 0 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
SARA (StAtutory Reasoning Assessment)
A dataset for statutory reasoning in tax law entailment and question answering.
13 papers · 0 benchmarks
Who-did-What (Who did What)
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
A dataset that contains 25,017 reading comprehension style examples curated from an existing corpus of 115 website privacy policies.
12 papers · 0 benchmarks
CMRC 2017 (Chinese Machine Reading Comprehension 2017)
Contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set.
11 papers · 0 benchmarks
ReCAM (SemEval-2021 Task 4: Reading Comprehension of Abstract Meaning)
Tasks Our shared task has three subtasks.
11 papers · 1 benchmark
SberQuAD (Sberbank Question Answering Dataset)
A large scale analogue of Stanford SQuAD in the Russian language - is a valuable resource that has not been properly presented to the scientific community.
10 papers · 1 benchmark
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
RUSSE (Russian Words in Context (based on RUSSE))
WiC: The Word-in-Context Dataset A reliable benchmark for the evaluation of context-sensitive word embeddings.
8 papers · 1 benchmark
The Video-based Multimodal Summarization with Multimodal Output (VMSMO) corpus consists of 184,920 document-summary pairs, with 180,000 training pairs, 2,460 validation and test pairs.
8 papers · 0 benchmarks
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
Contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.
5 papers · 0 benchmarks
OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance.
5 papers · 0 benchmarks
Taiga Corpus (An open-source corpus for machine learning.)
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
UIT-ViNewsQA is a new corpus for the Vietnamese language to evaluate healthcare reading comprehension models.
5 papers · 0 benchmarks
CS (Chinese Simile)
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
3 papers · 0 benchmarks
ConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e.
3 papers · 1 benchmark
DMQA (DeepMind Q&A)
The DeepMind Q&A Dataset consists of two datasets for Question Answering, CNN and DailyMail.
3 papers · 0 benchmarks
Shmoop Corpus is a dataset of 231 stories that are paired with detailed multi-paragraph summaries for each individual chapter (7,234 chapters), where the summary is chronologically aligned with respect to the story chapter.
3 papers · 0 benchmarks
A dataset containing 2,221 questions from matriculation exams for twelfth grade in various subjects -history, biology, geography and philosophy-, and 412 additional questions from online quizzes in history.
2 papers · 0 benchmarks
CodeQueries Benchmark dataset consists of instances of semantic queries, code context and code spans in the context corresponding to the semantic queries.
2 papers · 0 benchmarks
A dataset of around 2 million examples for machine reading-comprehension.
2 papers · 0 benchmarks
A newly developed public dataset and the task of multiple property extraction.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
JECC (Jericho Environment Commonsense Comprehension)
Jericho Environment Commonsense Comprehension (JECC) is a dataset for commonsense reasoning.
1 paper · 0 benchmarks
A large-scale machine comprehension dataset (based on the COCO images and captions).
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
PersianQA (Persian Question Answering Dataset)
PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia.
1 paper · 0 benchmarks
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
1 paper · 0 benchmarks
To collect WikiSuggest, Google Suggest API is used to harvest natural language questions and submit them to Google Search.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.