Home › Datasets › task › Question Generation
Question Generation datasets
archive 2025-07-28
25 datasets carry the task tag "Question Generation" (the task itself: Question Generation), ordered by the archive's paper count. Page 1 of 1: 25 shown of 25. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Question Generation datasets 1–25 of 25
The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
The Natural Questions corpus is a question answering dataset containing 307,373 training examples, 7,830 development examples, and 7,842 test examples.
1,404 papers · 9 benchmarks
TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web.
953 papers · 5 benchmarks
OK-VQA (Outside Knowledge Visual Question Answering)
Outside Knowledge Visual Question Answering (OK-VQA) includes more than 14,000 questions that require external knowledge to answer.
368 papers · 2 benchmarks
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others.
183 papers · 1 benchmark
SQuAD (Stanford Question Answering Dataset)
The Stanford Question Answering Dataset (SQuAD) is a collection of question-answer pairs derived from Wikipedia articles.
151 papers · 12 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
A new large-scale question-answering dataset that requires reasoning on heterogeneous information.
70 papers · 1 benchmark
GrailQA (Strongly Generalizable Question Answering)
GrailQA is a new large-scale, high-quality dataset for question answering on knowledge bases (KBQA) on Freebase with 64,331 questions annotated with both answers and corresponding logical forms in different syntax (i.e., SPARQL,…
34 papers · 4 benchmarks
FairytaleQA is a dataset focusing on narrative comprehension of kindergarten to eighth-grade students.
26 papers · 2 benchmarks
ROPES (Reasoning Over Paragraph Effects in Situations)
ROPES is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation.
24 papers · 0 benchmarks
The MMD (MultiModal Dialogs) dataset is a dataset for multimodal domain-aware conversations.
18 papers · 0 benchmarks
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
A dataset of ~19K questions that are elicited while a person is reading through a document.
13 papers · 0 benchmarks
GLGE (General Language Generation Evaluation)
GLGE is a general language generation evaluation benchmark which is composed of 8 language generation tasks, including Abstractive Text Summarization (CNN/DailyMail, Gigaword, XSUM, MSNews), Answer-aware Question Generation (SQuAD 1.1,…
12 papers · 0 benchmarks
The question-answer (QA) pairs are automatically generated using state-of-the-art question generation methods based on paintings and comments provided in an existing art understanding dataset.
9 papers · 0 benchmarks
ARID (Autonomous Robot Indoor Dataset)
ARID is a large-scale, multi-view object dataset collected with an RGB-D camera mounted on a mobile robot.
5 papers · 0 benchmarks
DiSCQ (Discharge Summary Clinical Questions)
DiSCQ is a newly curated question dataset composed of 2,000+ questions paired with the snippets of text (triggers) that prompted each question.
4 papers · 0 benchmarks
ClarQ, consists of ∼2M examples distributed across 173 domains of stackexchange.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
Dataset Description The dataset described in the provided text is focused on social media polls collected from Weibo, a popular Chinese microblogging platform.
3 papers · 3 benchmarks
30MQA (30M Factoid Question-Answer Corpus)
An enormous question answer pair corpus produced by applying a novel neural network architecture on the knowledge base Freebase to transduce facts into natural language questions.
2 papers · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.