Home › Datasets › task › Open-Domain Question Answering
Open-Domain Question Answering datasets
archive 2025-07-28
26 datasets carry the task tag "Open-Domain Question Answering" (the task itself: Open-Domain Question Answering), ordered by the archive's paper count. Page 1 of 1: 26 shown of 26. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Open-Domain Question Answering datasets 1–26 of 26
The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
The WebQuestions dataset is a question answering dataset using Freebase as the knowledge base and contains 6,642 question-answer pairs.
241 papers · 4 benchmarks
LAMA (LAnguage Model Analysis)
LAnguage Model Analysis (LAMA) consists of a set of knowledge sources, each comprised of a set of facts.
214 papers · 0 benchmarks
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
SearchQA was built using an in-production, commercial search engine.
133 papers · 1 benchmark
KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks: Fact-checking (FEVER), Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB), Slot filling (T-Rex, Zero Shot RE), Open…
117 papers · 11 benchmarks
DuReader is a large-scale open-domain Chinese machine reading comprehension dataset.
65 papers · 1 benchmark
QUASAR-T (QUestion Answering by Search And Reading – Trivia)
QUASAR-T is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
55 papers · 1 benchmark
QReCC contains 14K conversations with 81K question-answer pairs.
51 papers · 0 benchmarks
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
TQA (Textbook Question Answering)
The TextbookQuestionAnswering (TQA) dataset is drawn from middle school science curricula.
48 papers · 1 benchmark
QUASAR (QUestion Answering by Search And Reading)
The Question Answering by Search And Reading (QUASAR) is a large-scale dataset consisting of QUASAR-S and QUASAR-T.
47 papers · 1 benchmark
Break is a question understanding dataset, aimed at training models to reason over complex questions.
39 papers · 0 benchmarks
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks
Composed of 1,395 questions posed by crowdworkers on Wikipedia articles, and a machine translation of the Stanford Question Answering Dataset (Arabic-SQuAD).
24 papers · 0 benchmarks
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
QAMPARI is an ODQA benchmark, where question answers are lists of entities, spread across many paragraphs.
19 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
ConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e.
3 papers · 1 benchmark
CREPE is QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.
2 papers · 0 benchmarks
CoreSearch is a dataset for Cross-Document Event Coreference Search.
1 paper · 0 benchmarks
WikiQAar (English-Arabic Wikipedia Question-Answering)
A publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.