Home › Datasets › task › Question Answering

Question Answering datasets

archive 2025-07-28

413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 3 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Question Answering datasets 97–144 of 413

Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
ReVerb Challenge (REverberant Voice Enhancement and Recognition Benchmark)
The REVERB (REverberant Voice Enhancement and Recognition Benchmark) challenge is a benchmark for evaluation of automatic speech recognition techniques.
55 papers · 1 benchmark
The Bamboogle dataset is a collection of questions that was constructed to investigate the ability of language models to perform compositional reasoning tasks.
54 papers · 1 benchmark
DAQUAR (DAtaset for QUestion Answering on Real-world images) is a dataset of human question answer pairs about images.
54 papers · 0 benchmarks
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark
QReCC contains 14K conversations with 81K question-answer pairs.
51 papers · 0 benchmarks
Representation and learning of commonsense knowledge is one of the foundational problems in the quest to enable deep language understanding.
51 papers · 1 benchmark
TGIF (Tumblr GIF)
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
TOFU (Task of Fictitious Unlearning)
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks.
51 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
50 papers · 1 benchmark
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
emrQA has 1 million question-logical form and 400,000+ questionanswer evidence pairs.
50 papers · 0 benchmarks
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
TQA (Textbook Question Answering)
The TextbookQuestionAnswering (TQA) dataset is drawn from middle school science curricula.
48 papers · 1 benchmark
QUASAR (QUestion Answering by Search And Reading)
The Question Answering by Search And Reading (QUASAR) is a large-scale dataset consisting of QUASAR-S and QUASAR-T.
47 papers · 1 benchmark
CoS-E (Commonsense Explanations Dataset)
CoS-E consists of human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations Source: Explain Yourself!
46 papers · 0 benchmarks
GeoQA (Geometric Question Answering)
GeoQA is a dataset for automatic geometric problem solving containing 5,010 geometric problems with corresponding annotated programs, which illustrate the solving process of the given problems Compared with another publicly available…
46 papers · 1 benchmark
DART is a large dataset for open-domain structured data record to text generation.
45 papers · 3 benchmarks
A new large-scale geometry problem-solving dataset - 3,002 multi-choice geometry problems - dense annotations in formal language for the diagrams and text - 27,213 annotated diagram logic forms (literals) - 6,293 annotated text logic forms…
45 papers · 1 benchmark
ShARC (Shaping Answers with Rules through Conversation)
ShARC is a Conversational Question Answering dataset focussing on question answering from texts containing rules.
43 papers · 0 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
Contains around 200K dialogs with a total of 1.6M turns.
40 papers · 0 benchmarks
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
Break is a question understanding dataset, aimed at training models to reason over complex questions.
39 papers · 0 benchmarks
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
ConvFinQA (Conversational Finance Question Answering)
ConvFinQA is a dataset designed to study the chain of numerical reasoning in conversational question answering.
38 papers · 2 benchmarks
InsuranceQA is a question answering dataset for the insurance domain, the data stemming from the website Insurance Library.
38 papers · 0 benchmarks
TyDiQA is the gold passage version of the Typologically Diverse Question Answering (TyDiWA) dataset, a benchmark for information-seeking question answering, which covers nine languages.
38 papers · 1 benchmark
Doc2Dial (Doc2Dial: Document-grounded Dialogue)
For goal-oriented document-grounded dialogs, it often involves complex contexts for identifying the most relevant information, which requires better understanding of the inter-relations between conversations and documents.
36 papers · 0 benchmarks
The MSLR-WEB10K dataset consists of 10,000 search queries over the documents from search results.
36 papers · 0 benchmarks
The Open Table-and-Text Question Answering (OTT-QA) dataset contains open questions which require retrieving tables and text from the web to answer.
36 papers · 1 benchmark
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
MeQSum is a dataset for medical question summarization.
33 papers · 1 benchmark
SCIREX is a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N-ary relation identification from scientific articles.
33 papers · 2 benchmarks
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
QuaRTz (QuaRTz Dataset)
QuaRTz is a crowdsourced dataset of 3864 multiple-choice questions about open domain qualitative relationships.
30 papers · 0 benchmarks
CODAH (COmmonsense Dataset Adversarially-authored by Humans)
The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense question-answering in the sentence completion style of SWAG.
29 papers · 2 benchmarks
A machine reading comprehension (MRC) dataset with discourse structure built over multiparty dialog.
29 papers · 2 benchmarks
KaggleDBQA (KaggleDBQA: Realistic Text-to-SQL dataset)
KaggleDBQA is a challenging cross-domain and complex evaluation dataset of real Web databases, with domain-specific data types, original formatting, and unrestricted questions.
28 papers · 1 benchmark
decaNLP (Natural Language Decathlon Benchmark)
Natural Language Decathlon Benchmark (decaNLP) is a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation…
28 papers · 0 benchmarks
CaseHOLD (Case Holdings On Legal Decisions)
CaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.
27 papers · 2 benchmarks
VQA-HAT (VQA Human Attention)
VQA-HAT (Human ATtention) is a dataset to evaluate the informative regions of an image depending on the question being asked about it.
27 papers · 0 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.