Browse State-of-the-Art › Response Generation
Response Generation
370 papers with code · 3 benchmarks · 8 datasets archive 2025-07-28
A task where an agent should play the DE role and generate a text to respond to a P message.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SIMMC2.0 (5 rows) | PaCE | PaCE: Unified Multi-modal Dialogue Pre-training with Progressive... | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
| ArgSciChat (3 rows) | LED(Q,F) | ArgSciChat: A Dataset for Argumentative Dialogues on Scientific Papers | code | — | Compare |
| MMConv (2 rows) | PaCE | PaCE: Unified Multi-modal Dialogue Pre-training with Progressive... | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 370 papers with code (914 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Feb 2019 21 repositories listedNatural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
-
11 Oct 2015 14 repositories listed Syntology ran 2 of 13 samples · 11 unverifiedSequence-to-sequence neural network models for generation of conversational responses tend to generate safe, commonplace responses (e.
-
8 May 2019 9 repositories listedThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks.
-
19 May 2016 9 repositories listedSequential data often possesses a hierarchical structure with complex dependencies between subsequences, such as found between the utterances in a dialogue.
-
15 May 2023 7 repositories listedOur proposed framework provides access: (i) for verifying whether automatic metrics are faithful to human preference, regardless of their correlation level to human; and (ii) for inspecting the strengths and limitations…
-
7 May 2019 7 repositories listedPre-training and fine-tuning, e.
-
17 Oct 2023 6 repositories listed Syntology ran 8 of 14 samples · 6 unverified · 3 pointer-only (licence)Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its own generations using special tokens, called reflection tokens.
-
1 Nov 2019 6 repositories listed Syntology ran 2 of 12 samples · 10 unverified · 12 pointer-only (licence)We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer).
-
29 Sep 2018 5 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedEven though machine learning has become the major scene in dialogue research community, the real breakthrough has been blocked by the scale of data available.
-
17 Jun 2024 4 repositories listed Syntology ran 20 of 30 samples · 10 unverified · 15 pointer-only (licence)Iterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones.
-
12 Sep 2019 4 repositories listedIn this work, we introduce the the Schema-Guided Dialogue (SGD) dataset, containing over 16k multi-domain conversations spanning 16 domains.
-
2 Jun 2016 4 repositories listedWe introduce the multiresolution recurrent neural network, which extends the sequence-to-sequence framework to model natural language generation as two parallel discrete stochastic processes: a sequence of high-level…
-
30 Mar 2023 3 repositories listedMotivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement.
-
30 Jun 2020 3 repositories listedTo build a high-quality open-domain chatbot, we introduce the effective training process of PLATO-2 via curriculum learning.
-
17 Oct 2019 3 repositories listedPre-training models have been proved effective for a wide range of natural language processing tasks.
-
11 Sep 2019 3 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedWe present CoSQL, a corpus for building cross-domain, general-purpose database (DB) querying dialogue systems.
-
19 Jun 2018 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedOpen domain response generation has achieved remarkable progress in recent years, but sometimes yields short and uninformative responses.
-
31 May 2018 3 repositories listedVariational autoencoders~(VAEs) have shown a promise in data-driven conversation modeling.
-
29 Jun 2017 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, previous work in dialogue response generation has shown that these metrics do not correlate strongly with human judgment in the non task-oriented dialogue setting.
-
14 Mar 2025 2 repositories listedIn response, we show evaluations of existing RAG methods which account for both context relevance and answer quality.
-
22 Dec 2024 2 repositories listedThis survey examines the state of the art in text summarization models, with a specific focus on the abstractive summarization approach.
-
11 Oct 2024 2 repositories listedReinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks.
-
13 Jun 2024 2 repositories listedSubsequent analyses of these findings unveil insights into the effectiveness of various types of commonsense in generating responses and the particular response traits enhanced through explicit reasoning for commonsense…
-
11 Oct 2023 2 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)In response to these challenges, we present an iterative search and reasoning framework, which consists of a textual encoder, a visual encoder, and a generator.
-
9 Oct 2023 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)This paper proposes a method called Selective Context that enhances the inference efficiency of LLMs by identifying and pruning redundancy in the input context to make the input more compact.
-
29 Sep 2023 2 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)Extensive experimental results demonstrate that expressive instructions are crucial to instruction-based image editing, and our MGIE can lead to a notable improvement in automatic metrics and human evaluation while…
-
26 Jul 2023 2 repositories listedHere we develop a novel few-shot overgenerate-and-rank approach that achieves the controlled generation of DAs.
-
4 May 2023 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedAs results of the experiment, our method shows competitive performance on the MultiWOZ benchmark compared to the existing end-to-end models.
-
20 Feb 2023 2 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedPrior works mainly focus on adopting advanced RL techniques to train the ToD agents, while the design of the reward function is not well studied.
-
13 Oct 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections