Home › Datasets › task › Question Answering
Question Answering datasets
archive 2025-07-28
413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 5 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Answering datasets 193–240 of 413
QED is a linguistically principled framework for explanations in question answering.
15 papers · 1 benchmark
A Benchmark for Robust Multi-Hop Spatial Reasoning in Texts
15 papers · 1 benchmark
A new dataset for the low-resource language as Vietnamese to evaluate MRC models.
15 papers · 0 benchmarks
AmazonQA consists of 923k questions, 3.6M answers and 14M reviews across 156k products.
14 papers · 0 benchmarks
GraphQuestions is a characteristic-rich dataset designed for factoid question answering.
14 papers · 2 benchmarks
With the same format as WikiHop, the MedHop dataset is based on research paper abstracts from PubMed, and the queries are about interactions between pairs of drugs.
14 papers · 0 benchmarks
ANTIQUE is a collection of 2,626 open-domain non-factoid questions from a diverse set of categories.
13 papers · 0 benchmarks
GooAQ is a large-scale dataset with a variety of answer types.
13 papers · 0 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
The KLEJ benchmark (Kompleksowa Lista Ewaluacji Językowych) is a set of nine evaluation tasks for the Polish language understanding task.
13 papers · 0 benchmarks
The first summarization collection containing question-driven summaries of answers to consumer health questions.
13 papers · 0 benchmarks
MedConceptsQA - Open Source Medical Concepts QA Benchmark The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA
13 papers · 2 benchmarks
Provides detailed, graph-based annotations of social situations depicted in movie clips.
13 papers · 0 benchmarks
SARA (StAtutory Reasoning Assessment)
A dataset for statutory reasoning in tax law entailment and question answering.
13 papers · 0 benchmarks
Super-CLEVR is a dataset for Visual Question Answering (VQA) where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently.
13 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
ConditionalQA is a Question Answering (QA) dataset that contains complex questions with conditional answers, i.e.
12 papers · 1 benchmark
The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence.
12 papers · 1 benchmark
A platform for research in embodied artificial intelligence (AI).
12 papers · 0 benchmarks
PROST (Physical Reasoning about Objects Through Space and Time)
The PROST (Physical Reasoning about Objects Through Space and Time) dataset contains 18,736 multiple-choice questions made from 14 manually curated templates, covering 10 physical reasoning concepts.
12 papers · 0 benchmarks
We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset.
12 papers · 3 benchmarks
A dataset that contains 25,017 reading comprehension style examples curated from an existing corpus of 115 website privacy policies.
12 papers · 0 benchmarks
A rich, extensible and efficient environment that contains 45,622 human-designed 3D scenes of visually realistic houses, ranging from single-room studios to multi-storied houses, equipped with a diverse set of fully labeled 3D objects,…
11 papers · 0 benchmarks
MULTITQ is a large-scale dataset featuring ample relevant facts and multiple temporal granularities.
11 papers · 1 benchmark
ProtoQA is a question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations.
11 papers · 0 benchmarks
Existing benchmarks for temporal QA focus on a single information source (either a KB or a text corpus), and include only few questions with implicit constraints.
11 papers · 1 benchmark
CLEVR-Dialog is a large diagnostic dataset for studying multi-round reasoning in visual dialog.
10 papers · 0 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
A challenging new benchmark for language-agnostic answer retrieval from a multilingual candidate pool.
10 papers · 0 benchmarks
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
ReQA (Retrieval Question-Answering)
Retrieval Question-Answering (ReQA) benchmark tests a model’s ability to retrieve relevant answers efficiently from a large set of documents.
10 papers · 0 benchmarks
SberQuAD (Sberbank Question Answering Dataset)
A large scale analogue of Stanford SQuAD in the Russian language - is a valuable resource that has not been properly presented to the scientific community.
10 papers · 1 benchmark
A corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities.
10 papers · 0 benchmarks
The question-answer (QA) pairs are automatically generated using state-of-the-art question generation methods based on paintings and comments provided in an existing art understanding dataset.
9 papers · 0 benchmarks
CQASUMM is a dataset for CQA (Community Question Answering) summarization, constructed from the 4.4 million Yahoo!
9 papers · 0 benchmarks
ComQA is a large dataset of real user questions that exhibit different challenging aspects such as compositionality, temporal reasoning, and comparisons.
9 papers · 0 benchmarks
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LEAF-QA, a comprehensive dataset of 250,000 densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and semantics of these charts.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
SPARTQA (SPAtial Reasoning on Textual Question Answering)
SpartQA is a textual question answering benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior datasets and that is challenging for state-of-the-art language models…
9 papers · 0 benchmarks
SPARTQA - (SPAtial Reasoning on Textual Question Answering.)
We take advantage of the ground truth of NLVR images, design CFGs to generate stories, and use spatial reasoning rules to ask and answer spatial reasoning questions.
9 papers · 0 benchmarks
Provides four new test sets for the Stanford Question Answering Dataset (SQuAD) and evaluate the ability of question-answering systems to generalize to new data.
9 papers · 3 benchmarks
SelQA is a dataset that consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English Wikipedia.
9 papers · 0 benchmarks
WebCPM is a Chinese LFQA dataset.
9 papers · 0 benchmarks
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.