Home › Datasets › task › Question Answering
Question Answering datasets
archive 2025-07-28
413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 4 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Answering datasets 145–192 of 413
DUDE (Document UnderstanDing of Everything)
DUDE is formulated as an instance of Document Question Answering (DocQA) to evaluate how well current solutions deal with multi-page documents, if they can navigate and reason over the layout, and if they can generalize these skills to…
26 papers · 0 benchmarks
FairytaleQA is a dataset focusing on narrative comprehension of kindergarten to eighth-grade students.
26 papers · 2 benchmarks
The GenericsKB contains 3.4M+ generic sentences about the world, i.e., sentences expressing general truths such as "Dogs bark," and "Trees remove carbon dioxide from the atmosphere." Generics are potentially useful as a knowledge source…
26 papers · 0 benchmarks
MultiDoc2Dial (MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents)
MultiDoc2Dial is a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.
26 papers · 0 benchmarks
WikiReading is a large-scale natural language understanding task and publicly-available dataset with 18 million instances.
26 papers · 0 benchmarks
ASNQ (Answer Sentence Natural Questions)
A large scale dataset to enable the transfer step, exploiting the Natural Questions dataset.
25 papers · 1 benchmark
DIOR-RSVG is a large-scale benchmark dataset of remote sensing data (RSVG).
25 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
Composed of 1,395 questions posed by crowdworkers on Wikipedia articles, and a machine translation of the Stanford Question Answering Dataset (Arabic-SQuAD).
24 papers · 0 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
In SpokenSQuAD, the document is in spoken form, the input question is in the form of text and the answer to each question is always a span in the document.
24 papers · 1 benchmark
Here, we take a key step in this direction and release a new benchmark, TempQuestions, containing 1,271 questions, that are all temporal in nature, paired with their answers.
24 papers · 1 benchmark
HeadQA is a multi-choice question answering testbed to encourage research on complex reasoning.
23 papers · 1 benchmark
A large-scale dataset for Complex KBQA.
23 papers · 1 benchmark
SituatedQA is an open-retrieval QA dataset where systems must produce the correct answer to a question given the temporal or geographical context.
23 papers · 0 benchmarks
Social-IQ is an unconstrained benchmark specifically designed to train and evaluate socially intelligent technologies.
23 papers · 0 benchmarks
Question answering over knowledge graphs (KG-QA) is a vital topic in IR.
23 papers · 1 benchmark
BookTest is a new dataset similar to the popular Children’s Book Test (CBT), however more than 60 times larger.
22 papers · 0 benchmarks
A new benchmark dataset for cross-lingual and multilingual question answering for high school examinations.
22 papers · 0 benchmarks
TopiOCQA (pronounced Tapioca) is an open-domain conversational dataset with topic switches on Wikipedia.
22 papers · 0 benchmarks
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
FQuAD (French Question Answering Dataset)
A French Native Reading Comprehension dataset of questions and answers on a set of Wikipedia articles that consists of 25,000+ samples for the 1.0 version and 60,000+ samples for the 1.1 version.
21 papers · 1 benchmark
WIQA (What-If Question Answering)
The WIQA dataset V1 has 39705 questions containing a perturbation and a possible effect in the context of a paragraph.
21 papers · 0 benchmarks
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
MINTAKA is a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models.
20 papers · 0 benchmarks
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
20 papers · 0 benchmarks
One of the largest commonsense knowledge bases available, describing over 2 million disambiguated concepts and activities, connected by over 18 million assertions.
20 papers · 0 benchmarks
CRONQUESTIONS, the Temporal KGQA dataset consists of two parts: a KG with temporal annotations, and a set of natural language questions requiring temporal reasoning.
19 papers · 1 benchmark
This dataset code generates mathematical question and answer pairs, from a range of question types at roughly school-level difficulty.
19 papers · 1 benchmark
QAMPARI is an ODQA benchmark, where question answers are lists of entities, spread across many paragraphs.
19 papers · 0 benchmarks
A dataset on asking Questions for Lack of Clarity in open-domain information-seeking conversations.
19 papers · 0 benchmarks
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
StaQC (Stack Overflow Question-Code pairs) is a large dataset of around 148K Python and 120K SQL domain question-code pairs, which are automatically mined from StackOverflow.
19 papers · 0 benchmarks
Torque is an English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.
19 papers · 1 benchmark
With social media becoming increasingly popular on which lots of news and real-time events are reported, developing automated question answering systems is critical to the effectiveness of many applications that rely on real-time knowledge.
19 papers · 1 benchmark
A dataset with 2,437 dialogues and 10,917 QA pairs.
18 papers · 0 benchmarks
ORCAS is a click-based dataset.
18 papers · 0 benchmarks
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
ConvQuestions is the first realistic benchmark for conversational question answering over knowledge graphs.
17 papers · 0 benchmarks
EgoTask QA benchmark contains 40K balanced question-answer pairs selected from 368K programmatically generated questions generated over 2K egocentric videos.
17 papers · 1 benchmark
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
ORConvQA (Open-Retrieval Conversational Question Answering)
Enhances QuAC by adapting it to an open-retrieval setting.
17 papers · 0 benchmarks
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
ToolQA is a question answering benchmark for Large Language Models (LLMs) which is designed to faithfully evaluate LLMs' ability to use external tools for question answering.
16 papers · 0 benchmarks
The beginnings of a question answering dataset specifically designed for COVID-19, built by hand from knowledge gathered from Kaggle's COVID-19 Open Research Dataset Challenge.
15 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.