Browse State-of-the-Art › Question Generation
Question Generation
264 papers with code · 11 benchmarks · 25 datasets archive 2025-07-28
The goal of Question Generation is to generate a valid and fluent question according to a given passage and the target answer. Question Generation can be used in many scenarios, such as automatic tutoring systems, improving the performance of Question Answering models and enabling chatbots to lead a conversation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
11 leaderboard tables shown for this task, 11 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 11 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
25 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 264 papers with code (664 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Oct 2019 57 repositories listed Syntology ran 2 of 31 samples · 29 unverifiedTransfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP).
-
7 Oct 2016 25 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 14 pointer-only (licence)We observe that our method consistently outperforms BS and previously proposed techniques for diverse decoding from neural sequence models.
-
7 Oct 2016 25 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 14 pointer-only (licence)We observe that our method consistently outperforms BS and previously proposed techniques for diverse decoding from neural sequence models.
-
29 Apr 2017 10 repositories listedWe study automatic question generation for sentences from text passages in reading comprehension.
-
29 Apr 2017 10 repositories listedWe study automatic question generation for sentences from text passages in reading comprehension.
-
8 May 2019 9 repositories listedThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks.
-
8 May 2019 9 repositories listedThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks.
-
6 Apr 2017 6 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Automatic question generation aims to generate questions from a text passage where the generated questions can be answered by certain sub-spans of the given passage.
-
6 Apr 2017 6 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Automatic question generation aims to generate questions from a text passage where the generated questions can be answered by certain sub-spans of the given passage.
-
26 Jan 2020 5 repositories listedCurrent pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks.
-
26 Jan 2020 5 repositories listedCurrent pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks.
-
13 Jan 2020 5 repositories listedThis paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism.
-
13 Jan 2020 5 repositories listedThis paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism.
-
24 Oct 2024 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedDespite the availability of several open-source multimodal datasets, limitations in the scale and quality of open-source instruction data hinder the performance of VLMs trained on these datasets, leading to a…
-
24 Oct 2024 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedDespite the availability of several open-source multimodal datasets, limitations in the scale and quality of open-source instruction data hinder the performance of VLMs trained on these datasets, leading to a…
-
23 Dec 2020 4 repositories listedOpen-domain question answering can be reformulated as a phrase retrieval problem, without the need for processing documents on-demand during inference (Seo et al., 2019).
-
3 May 2020 4 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedQuestion generation (QG) is a natural language generation task where a model is trained to ask questions corresponding to some input text.
-
3 May 2020 4 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedQuestion generation (QG) is a natural language generation task where a model is trained to ask questions corresponding to some input text.
-
12 Jun 2019 4 repositories listedWe introduce a novel method of generating synthetic question answering corpora by combining models of question generation and answer extraction, and by filtering the results to ensure roundtrip consistency.
-
12 Jun 2019 4 repositories listedWe introduce a novel method of generating synthetic question answering corpora by combining models of question generation and answer extraction, and by filtering the results to ensure roundtrip consistency.
-
4 May 2017 4 repositories listedWe propose a recurrent neural model that generates natural-language questions from documents, conditioned on answers.
-
4 May 2017 4 repositories listedWe propose a recurrent neural model that generates natural-language questions from documents, conditioned on answers.
-
7 Mar 2022 3 repositories listedWe introduce IT5, the first family of encoder-decoder transformer models pretrained specifically on Italian.
-
8 Sep 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWe crafted questions that some humans would answer falsely due to a false belief or misconception.
-
16 Apr 2021 3 repositories listedProphetNet is a pre-training based natural language generation method which shows powerful performance on English text summarization and question generation tasks.
-
16 Apr 2021 3 repositories listedProphetNet is a pre-training based natural language generation method which shows powerful performance on English text summarization and question generation tasks.
-
1 Nov 2020 3 repositories listedThis paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism.
-
1 Nov 2020 3 repositories listedThis paper presents a new sequence-to-sequence pre-training model called ProphetNet, which introduces a novel self-supervised objective named future n-gram prediction and the proposed n-stream self-attention mechanism.
-
16 Jul 2020 3 repositories listed Syntology ran 2 of 12 samples · 10 unverifiedWe show that the PLMs BART and T5 achieve new state-of-the-art results and that task-adaptive pretraining strategies improve their performance even further.
-
28 Feb 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM).
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections