Home › Datasets › task › Question Answering

Question Answering datasets

archive 2025-07-28

413 datasets carry the task tag "Question Answering" (the task itself: Question Answering), ordered by the archive's paper count. Page 6 of 9: 48 shown of 413. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Question Answering datasets 241–288 of 413

GermanQuAD is a Question Answering (QA) dataset of 13,722 extractive question/answer pairs in German.
8 papers · 1 benchmark
OPIEC (Open Information Extraction Corpus)
OPIEC is an Open Information Extraction (OIE) corpus, constructed from the entire English Wikipedia.
8 papers · 0 benchmarks
Peyma is a Persian NER dataset to train and test NER systems.
8 papers · 0 benchmarks
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
Consists of multiple sentences whose clues are arranged by difficulty (from obscure to obvious) and uniquely identify a well-known entity such as those found on Wikipedia.
8 papers · 1 benchmark
SubjQA is a question answering dataset that focuses on subjective (as opposed to factual) questions and answers.
8 papers · 0 benchmarks
A configurable visual question and answer dataset (COG) to parallel experiments in humans and animals.
7 papers · 0 benchmarks
DBLP-QuAD (DBLP Question Answering Dataset)
In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG).
7 papers · 0 benchmarks
DaNetQA (Yes/no Question Answering Dataset for the Russian)
DaNetQA is a question answering dataset for yes/no questions.
7 papers · 1 benchmark
ForecastQA is a question-answering dataset consisting of 10,392 event forecasting questions, which have been collected and verified via crowdsourcing efforts.
7 papers · 0 benchmarks
GHOSTS is the first natural-language dataset made and curated by working researchers in mathematics that (1) aims to cover graduate-level mathematics and (2) provides a holistic overview of the mathematical capabilities of language models.
7 papers · 0 benchmarks
IQUAD (Interactive Question Answering Dataset)
IQUAD is a dataset for Visual Question Answering in interactive environments.
7 papers · 0 benchmarks
JGLUE, Japanese General Language Understanding Evaluation, is built to measure the general NLU ability in Japanese.
7 papers · 0 benchmarks
A new question answering dataset constructed from play-by-play live broadcast.
7 papers · 0 benchmarks
Collects the data by scraping Wikipedia and then utilize crowdsourcing to collect question-answer pairs.
7 papers · 0 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
7 papers · 1 benchmark
RuBQ (Russian Knowledge Base Questions)
The first Russian knowledge base question answering (KBQA) dataset.
7 papers · 0 benchmarks
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
k-qa (K-QA: A Real-World Medical Q&A Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 0 benchmarks
BIPIA (Benchmark of Indirect Prompt Injection Attacks)
Recent advancements in large language models (LLMs) have led to their adoption across various applications, notably in combining LLMs with external content to generate responses.
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
KAMEL (Knowledge Analysis with Multitoken Entities in Language Models)
KAMEL comprises knowledge about 234 relations from Wikidata with a large training, validation, and test dataset.
6 papers · 1 benchmark
MovieFIB (Movie Fill-in-the-Blank)
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
V2C (Video-to-Commonsense)
6 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
AI2D-RST is a multimodal corpus of 1000 English-language diagrams that represent topics in primary school natural sciences, such as food webs, life cycles, moon phases and human physiology.
5 papers · 0 benchmarks
A synthetically generated QA dataset for text-based reasoning.
5 papers · 0 benchmarks
KG20C (A scholarly knowledge graph benchmark dataset)
KG20C is a Knowledge Graph about high quality papers from 20 top computer science Conferences.
5 papers · 1 benchmark
MATINF (Maternal and Infant Dataset)
Maternal and Infant (MATINF) Dataset is a large-scale dataset jointly labeled for classification, question answering and summarization in the domain of maternity and baby caring in Chinese.
5 papers · 0 benchmarks
Contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.
5 papers · 0 benchmarks
MoA (MoA_Long_ModelQA)
This is the dataset used by the automatic sparse attention compression method MoA.
5 papers · 0 benchmarks
OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance.
5 papers · 0 benchmarks
PQuAD (Persian Question Answering Dataset)
Persian Question Answering Dataset (PQuAD) is a crowdsourced reading comprehension dataset on Persian Wikipedia articles.
5 papers · 0 benchmarks
QAConv is a new question answering (QA) dataset that uses conversations as a knowledge source.
5 papers · 0 benchmarks
RuCoS (Russian Reading Comprehension with Commonsense Reasoning)
Russian reading comprehension with Commonsense reasoning (RuCoS) is a large-scale reading comprehension dataset that requires commonsense reasoning.
5 papers · 1 benchmark
SpaRTUN a dataset synthesized for transfer learning on spatial question answering (SQA) and spatial role labeling (SpRL).
5 papers · 0 benchmarks
TutorialVQA is a new type of dataset used to find answer spans in tutorial videos.
5 papers · 0 benchmarks
WikiHowQA is a Community-based Question Answering dataset, which can be used for both answer selection and abstractive summarization tasks.
5 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CS1QA is a dataset for code-based question answering in the programming education domain.
4 papers · 0 benchmarks
CompMix is a crowdsourced QA benchmark which naturally demands the integration of a mixture of input sources.
4 papers · 0 benchmarks
ConvRef is a conversational QA benchmark with reformulations.
4 papers · 0 benchmarks
DiSCQ (Discharge Summary Clinical Questions)
DiSCQ is a newly curated question dataset composed of 2,000+ questions paired with the snippets of text (triggers) that prompted each question.
4 papers · 0 benchmarks
IndoNLG is a benchmark to measure natural language generation (NLG) progress in three low-resource—yet widely spoken—languages of Indonesia: Indonesian, Javanese, and Sundanese.
4 papers · 0 benchmarks
JaQuAD (Japanese Question Answering Dataset) is a question answering dataset in Japanese that consists of 39,696 extractive question-answer pairs on Japanese Wikipedia articles.
4 papers · 1 benchmark
SCDE is a human-created sentence cloze dataset, collected from public school English examinations in China.
4 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.