Browse State-of-the-Art › Dialogue Generation
Dialogue Generation
265 papers with code · 13 benchmarks · 34 datasets archive 2025-07-28
Dialogue generation is the task of "understanding" natural language inputs - within natural language processing in order to produce output. The systems are usually intended for conversing with humans, for instance back and forth dialogue with a conversation agent like a chatbot. Some example benchmarks for this task (see others such as Natural Language Understanding) include FusedChat and Ubuntu DIalogue Corpus (UDC). Models can be evaluated via metrics such as BLEU, ROUGE, and METEOR albeit with challenges in terms of weak correlation with human judgement, that may be addressed by new ones like UnSupervised and Reference-free (USR) and Metric for automatic Unreferenced dialog evaluation (MaUde).
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
13 leaderboard tables shown for this task, 13 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 13 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
34 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 34 until expanded.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 265 papers with code (606 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 Sep 2014 124 repositories listed Syntology ran 21 of 44 samples · 23 unverified · 16 pointer-only (licence)Neural machine translation is a recently proposed approach to machine translation.
-
23 Jan 2019 23 repositories listed Syntology ran 8 of 26 samples · 18 unverified · 7 pointer-only (licence)We introduce a new approach to generative data-driven dialogue systems (e.
-
22 Jan 2018 15 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Chit-chat models are known to have several problems: they lack specificity, do not display a consistent personality and are often not very captivating.
-
1 Nov 2018 9 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)One challenge for dialogue agents is recognizing feelings in the conversation partner and replying accordingly, a key communicative skill.
-
5 Oct 2018 8 repositories listedWe propose several strong multimodal baselines and show the importance of contextual and multimodal information for emotion recognition in conversations.
-
23 Jan 2017 8 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)In this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from…
-
5 Jun 2016 8 repositories listedRecent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on…
-
15 May 2023 7 repositories listedOur proposed framework provides access: (i) for verifying whether automatic metrics are faithful to human preference, regardless of their correlation level to human; and (ii) for inspecting the strengths and limitations…
-
26 Apr 2021 5 repositories listedTo enhance the generalization ability of PanGu-α, we collect 1.
-
26 Jan 2020 5 repositories listedCurrent pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks.
-
8 Jul 2019 4 repositories listedMultiple sequence to sequence models were used to establish an end-to-end multi-turns proactive dialogue generation agent, with the aid of data augmentation techniques and variant encoder-decoder structure designs.
-
2 Jun 2016 4 repositories listedWe introduce the multiresolution recurrent neural network, which extends the sequence-to-sequence framework to model natural language generation as two parallel discrete stochastic processes: a sequence of high-level…
-
29 Mar 2023 3 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedIn this work, we present G-Eval, a framework of using large language models with chain-of-thoughts (CoT) and a form-filling paradigm, to assess the quality of NLG outputs.
-
20 Sep 2021 3 repositories listedTo explore the limit of dialogue generation pre-training, we present the models of PLATO-XL with up to 11 billion parameters, trained on both Chinese and English social media conversations.
-
17 Oct 2019 3 repositories listedPre-training models have been proved effective for a wide range of natural language processing tasks.
-
4 Apr 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Generating texts which express complex ideas spanning multiple sentences requires a structured representation of their content (document plan), but these representations are prohibitively expensive to manually produce.
-
23 Feb 2019 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedDefining action spaces for conversational agents and optimizing their decision-making process with reinforcement learning is an enduring challenge.
-
28 Jan 2019 3 repositories listedIn this paper, we investigate the problem of incorporating explicit personality traits in dialogue generation to deliver personalized dialogues.
-
3 Nov 2018 3 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn open-domain dialogue intelligent agents should exhibit the use of knowledge, however there are few convincing demonstrations of this to date.
-
5 Feb 2018 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Existing text generation methods tend to produce repeated and "boring" expressions.
-
28 Nov 2017 3 repositories listedThis paper presents a new adversarial learning method for generative conversational agents (GCA) besides a new model of GCA.
-
29 Jun 2017 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, previous work in dialogue response generation has shown that these metrics do not correlate strongly with human judgment in the non task-oriented dialogue setting.
-
29 Sep 2022 2 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedOpen-domain dialogue systems aim to interact with humans through natural language texts in an open-ended fashion.
-
23 May 2022 2 repositories listedThis work presents BanglaNLG, a comprehensive benchmark for evaluating natural language generation (NLG) models in Bangla, a widely spoken yet low-resource language.
-
31 Mar 2022 2 repositories listedWe investigate different aspects of responses generated by PanGu-Bot, including response quality, knowledge, and safety.
-
16 Sep 2020 2 repositories listedNeural dialogue response generation has gained much popularity in recent years.
-
10 Aug 2020 2 repositories listedThe cleaned dataset and the pre-training models will facilitate the research of short-text conversation modeling.
-
1 May 2020 2 repositories listedWe present a general approach towards controllable societal biases in natural language generation (NLG).
-
6 Feb 2020 2 repositories listedTo address this, we propose a neural topical expansion framework, namely Persona Exploration and Exploitation (PEE), which is able to extend the predefined user persona description with semantically correlated content…
-
12 Nov 2019 2 repositories listed Syntology ran 2 of 12 samples · 10 unverifiedFurther, to incorporate the target persona in the decoding process and to balance its contribution, an attention routing structure is devised in the decoder to merge features extracted from the target persona and…
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections