Home › Datasets › task › Question Answering
Question Answering datasets
archive 2025-07-28
413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 2 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Answering datasets 49–96 of 413
MathQA significantly enhances the AQuA dataset with fully-specified operational programs.
159 papers · 1 benchmark
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
TyDiQA (Typologically Diverse Question Answering)
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 200K question-answer pairs.
148 papers · 0 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
e-SNLI is used for various goals, such as obtaining full sentence justifications of a model's decisions, improving universal sentence representations and transferring to out-of-domain NLI datasets.
139 papers · 1 benchmark
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
SIQA (Social Interaction QA)
Social Interaction QA (SIQA) is a question-answering benchmark for testing social commonsense intelligence.
120 papers · 1 benchmark
SimpleQuestions is a large-scale factoid question answering dataset.
120 papers · 2 benchmarks
KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks: Fact-checking (FEVER), Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB), Slot filling (T-Rex, Zero Shot RE), Open…
117 papers · 11 benchmarks
A dataset of large scale alignments between Wikipedia abstracts and Wikidata triples.
117 papers · 1 benchmark
The MRQA (Machine Reading for Question Answering) dataset is a dataset for evaluating the generalization capabilities of reading comprehension systems.
116 papers · 1 benchmark
MCTest is a freely available set of stories and associated questions intended for research on the machine comprehension of text.
114 papers · 2 benchmarks
QASC (Question Answering via Sentence Composition)
QASC is a question-answering dataset with a focus on sentence composition.
114 papers · 0 benchmarks
FinQA is a new large-scale dataset with Question-Answering pairs over Financial reports, written by financial experts.
110 papers · 1 benchmark
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
Uses structured and unstructured data.
102 papers · 0 benchmarks
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
QASPER is a dataset for question answering on scientific research papers.
102 papers · 1 benchmark
QuALITY (Question Answering with Long Input Texts, Yes!)
QuALITY (Question Answering with Long Input Texts, Yes!) is a multiple-choice question answering dataset for long document comprehension.
98 papers · 1 benchmark
RAVEN consists of 1,120,000 images and 70,000 RPM (Raven's Progressive Matrices) problems, equally distributed in 7 distinct figure configurations.
96 papers · 0 benchmarks
CBT (Children’s Book Test)
Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context.
92 papers · 1 benchmark
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
PopQA is an open-domain QA dataset with 14k QA pairs with fine-grained Wikidata entity ID, Wikipedia page views, and relationship type information.
91 papers · 1 benchmark
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
Logical reasoning is an important ability to examine, analyze, and critically evaluate arguments as they occur in ordinary language as the definition from Law School Admission Council.
88 papers · 4 benchmarks
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
WikiTableQuestions is a question answering dataset over semi-structured tables.
79 papers · 2 benchmarks
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
76 papers · 1 benchmark
The Semantic Scholar corpus (S2) is composed of titles from scientific papers published in machine learning conferences and journals from 1985 to 2017, split by year (33 timesteps).
75 papers · 0 benchmarks
TrecQA (Text Retrieval Conference Question Answering)
Text Retrieval Conference Question Answering (TrecQA) is a dataset created from the TREC-8 (1999) to TREC-13 (2004) Question Answering tracks.
73 papers · 3 benchmarks
A new large-scale question-answering dataset that requires reasoning on heterogeneous information.
70 papers · 1 benchmark
BeaverTails is a dataset aimed at fostering research on safety alignment in large language models (LLMs).
68 papers · 0 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
DREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset.
68 papers · 2 benchmarks
CFQ (Compositional Freebase Questions)
A large and realistic natural language question answering dataset.
66 papers · 1 benchmark
WikiHop is a multi-hop question-answering dataset.
66 papers · 2 benchmarks
COCO-QA is a dataset for visual question answering.
62 papers · 0 benchmarks
ComplexWebQuestions is a dataset for answering complex questions that require reasoning over multiple web snippets.
62 papers · 2 benchmarks
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
The WebQuestionsSP dataset is released as part of our ACL-2016 paper “The Value of Semantic Parse Labeling for Knowledge Base Question Answering” [Yih, Richardson, Meek, Chang & Suh, 2016], in which we evaluated the value of gathering…
61 papers · 3 benchmarks
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
CANARD (A Dataset for Question-in-Context Rewriting)
CANARD is a dataset for question-in-context rewriting that consists of questions each given in a dialog context together with a context-independent rewriting of the question.
58 papers · 1 benchmark
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
PrOntoQA (Proof and Ontology-Generated Question-Answering)
PrOntoQA is a question-answering dataset which generates examples with chains-of-thought that describe the reasoning required to answer the questions correctly.
55 papers · 0 benchmarks
QUASAR-T (QUestion Answering by Search And Reading – Trivia)
QUASAR-T is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
55 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.