Home › Datasets › task › Question Answering
Question Answering datasets
archive 2025-07-28
413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 7 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Answering datasets 289–336 of 413
UniProtQA consists of proteins and textual queries about their functions and properties.
4 papers · 1 benchmark
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
A large-scale dataset built on questions from TyDi QA lacking same-language answers.
4 papers · 0 benchmarks
AfriQA is a cross-lingual QA dataset with a focus on African languages.
3 papers · 0 benchmarks
A comprehensive multi-task benchmark for the Polish language understanding, accompanied by an online leaderboard.
3 papers · 0 benchmarks
CHQ-Summ (Consumer Healthcare Question Summarization)
Contains 1507 domain-expert annotated consumer health questions and corresponding summaries.
3 papers · 0 benchmarks
COVID-Q consists of COVID-19 questions which have been annotated into a broad category (e.g.
3 papers · 0 benchmarks
ClarQ, consists of ∼2M examples distributed across 173 domains of stackexchange.
3 papers · 0 benchmarks
A filtered version of CronQuestions and which can better demonstrate the model’s inference ability for complex temporal questions.
3 papers · 1 benchmark
ConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e.
3 papers · 1 benchmark
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages.
3 papers · 0 benchmarks
Contains over 70,000 question-answer pairs from both structured tables and unstructured notes from a publicly available Electronic Health Record (EHR).
3 papers · 0 benchmarks
This is a medical multiple-choice dataset with explanations which can be used to interpret the answer.
3 papers · 0 benchmarks
FanOutQA is a high quality, multi-hop, multi-document benchmark for large language models using English Wikipedia as its knowledge base.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
Contains 25,165 textual news articles collected from hundreds of news media sites (e.g., Yahoo News, Google News, CNN News.) and 76,516 image posts shared on Flickr social media, which are annotated according to 412 real-world events.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
MultiQ is a multi-hop QA dataset for Russian, suitable for general open-domain question answering, information retrieval, and reading comprehension tasks.
3 papers · 1 benchmark
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
PubChemQA consists of molecules and their corresponding textual descriptions from PubChem.
3 papers · 1 benchmark
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
ResQ (Real-world Spatial Question Answering)
ReSQ is a real-world Spatial Question Answering dataset with human-generated questions built on an existing corpus with SpRL annotations.
3 papers · 0 benchmarks
Researchy Questions is a set of about 100k Bing queries that users spent the most effort on.
3 papers · 0 benchmarks
Shmoop Corpus is a dataset of 231 stories that are paired with detailed multi-paragraph summaries for each individual chapter (7,234 chapters), where the summary is chronologically aligned with respect to the story chapter.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
The dataset comprises 4,500 question-answer pairs collected from trusted medical sources, with at least one answer and at most four unique paraphrased answers per question
3 papers · 0 benchmarks
We present the AWS documentation corpus, an open-book QA dataset, which contains 25,175 documents along with 100 matched questions and answers.
2 papers · 0 benchmarks
Almawave-SLU is the first Italian dataset for Spoken Language Understanding (SLU).
2 papers · 0 benchmarks
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
MultiReQA is a cross-domain evaluation for retrieval question answering models.
2 papers · 0 benchmarks
NQuAD (Nuclear Question Answering Dataset)
NQuAD is a Nuclear Question Answering Dataset, which contains 700+ nuclear Question Answer pairs developed and verified by expert nuclear researchers.
2 papers · 0 benchmarks
PoseScript is a dataset that pairs a few thousand 3D human poses from AMASS with rich human-annotated descriptions of the body parts and their spatial relationships.
2 papers · 0 benchmarks
QUITE (Quantifying Uncertainty in natural language Text)
QUITE (Quantifying Uncertainty in natural language Text) is an entirely new benchmark that allows for assessing the capabilities of neural language model-based systems w.r.t.
2 papers · 0 benchmarks
ReviewQA is a question-answering dataset based on hotel reviews.
2 papers · 0 benchmarks
RoMQA is a benchmark for robust, multi-evidence, and multi-answer question answering (QA).
2 papers · 0 benchmarks
RuOpenBookQA is a QA dataset with multiple-choice elementary-level science questions which probe the understanding of core science facts.
2 papers · 1 benchmark
Recent applications of LLMs in Machine Reading Comprehension (MRC) systems have shown impressive results, but the use of shortcuts, mechanisms triggered by features spuriously correlated to the true label, has emerged as a potential threat…
2 papers · 0 benchmarks
A new benchmark dataset for simple question answering over knowledge graphs that was created by mapping SimpleQuestions entities and predicates from Freebase to DBpedia.
2 papers · 0 benchmarks
Schema2QA is the first large question answering dataset over real-world Schema.org data.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.