Home › Datasets › task › Question Answering
Question Answering datasets
archive 2025-07-28
413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 1 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Answering datasets 1–48 of 413
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
MML (Massive Multitask Language Understanding)
MMLU (Massive Multitask Language Understanding) is a new benchmark designed to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.
1,922 papers · 29 benchmarks
The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
MS MARCO (Microsoft Machine Reading Comprehension Dataset)
The MS MARCO (Microsoft MAchine Reading Comprehension) is a collection of datasets focused on deep learning in search.
1,036 papers · 7 benchmarks
HellaSwag is a challenge dataset for evaluating commonsense NLI that is specially hard for state-of-the-art models, though its questions are trivial for humans (>95% accuracy).
994 papers · 6 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
HotpotQA is a question answering dataset collected on the English Wikipedia, containing about 113K crowd-sourced questions that are constructed to require the introduction paragraphs of two Wikipedia articles to answer.
933 papers · 3 benchmarks
ConceptNet is a knowledge graph that connects words and phrases of natural language with labeled edges.
852 papers · 1 benchmark
PIQA (Physical Interaction: Question Answering)
PIQA is a dataset for commonsense reasoning, and was created to investigate the physical knowledge of existing models in NLP.
772 papers · 2 benchmarks
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
BoolQ (Boolean Questions)
BoolQ is a question answering dataset for yes/no questions containing 15942 examples.
701 papers · 5 benchmarks
The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject.
635 papers · 3 benchmarks
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions.
607 papers · 3 benchmarks
CNN/Daily Mail is a dataset for text summarization.
530 papers · 8 benchmarks
FEVER (Fact Extraction and VERification)
FEVER is a publicly available dataset for fact extraction and verification against textual sources.
498 papers · 3 benchmarks
TextVQA is a dataset to benchmark visual reasoning based on text in images.
476 papers · 3 benchmarks
RACE (ReAding Comprehension dataset from Examinations)
The ReAding Comprehension dataset from Examinations (RACE) dataset is a machine reading comprehension dataset consisting of 27,933 passages and 97,867 questions from English exams, targeting Chinese students aged 12-18.
412 papers · 3 benchmarks
DROP (Discrete Reasoning Over Paragraphs)
Discrete Reasoning Over Paragraphs DROP is a crowdsourced, adversarially-created, 96k-question benchmark, in which a system must resolve references in a question, perhaps to multiple input positions, and perform discrete operations over…
382 papers · 3 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
BIG-Bench Hard (BBH) is a subset of the BIG-Bench, a diverse evaluation suite for language models.
352 papers · 2 benchmarks
Automatic image captioning is the task of producing a natural-language utterance (usually a sentence) that correctly reflects the visual content of an image.
352 papers · 2 benchmarks
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
COPA (Choice of Plausible Alternatives)
The Choice Of Plausible Alternatives (COPA) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
329 papers · 1 benchmark
Multiple choice question answering based on the United States Medical License Exams (USMLE).
328 papers · 1 benchmark
BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks.
311 papers · 10 benchmarks
StrategyQA is a question answering benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy.
291 papers · 1 benchmark
DocVQA consists of 50,000 questions defined on 12,000+ document images.
290 papers · 3 benchmarks
CoQA (Conversational Question Answering Challenge)
CoQA is a large-scale dataset for building Conversational Question Answering systems.
281 papers · 2 benchmarks
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
The NewsQA dataset is a crowd-sourced machine reading comprehension dataset of 120,000 question-answer pairs.
272 papers · 1 benchmark
WikiSQL consists of a corpus of 87,726 hand-annotated SQL query and natural language question pairs.
267 papers · 4 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
The WebQuestions dataset is a question answering dataset using Freebase as the knowledge base and contains 6,642 question-answer pairs.
241 papers · 4 benchmarks
LAMA (LAnguage Model Analysis)
LAnguage Model Analysis (LAMA) consists of a set of knowledge sources, each comprised of a set of facts.
214 papers · 0 benchmarks
AI2 Diagrams (AI2D) is a dataset of over 5000 grade school science diagrams with over 150000 rich annotations, their ground truth syntactic parses, and more than 15000 corresponding multiple choice questions.
207 papers · 1 benchmark
The NarrativeQA dataset includes a list of documents with Wikipedia summaries, links to full stories, and questions and answers.
206 papers · 1 benchmark
WikiQA (Wikipedia open-domain Question Answering)
The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
196 papers · 2 benchmarks
BioASQ (Biomedical Semantic Indexing and Question Answering)
BioASQ is a question answering dataset.
192 papers · 1 benchmark
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
QuAC (Question Answering in Context)
Question Answering in Context is a large-scale dataset that consists of around 14K crowdsourced Question Answering dialogs with 98K question-answer pairs in total.
178 papers · 1 benchmark
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
SWAG (Situations With Adversarial Generations)
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").
163 papers · 2 benchmarks
MultiRC (Multi-Sentence Reading Comprehension)
MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph.
162 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.