Browse State-of-the-Art › Open-Domain Dialog
Open-Domain Dialog
32 papers with code · 1 benchmark · 13 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| KILT: Wizard of Wikipedia (21 rows) | Hindsight | — | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
13 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 32 papers with code (60 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Apr 2021 3 repositories listedProphetNet is a pre-training based natural language generation method which shows powerful performance on English text summarization and question generation tasks.
-
4 Sep 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We test both task-specific and general baselines, evaluating downstream performance in addition to the ability of the models to provide provenance.
-
10 Mar 2021 2 repositories listedThe task of long-form question answering (LFQA) involves retrieving documents relevant to a given question and using them to generate a paragraph-length answer.
-
15 Sep 2020 2 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedParticularly, our ranker outperforms the conventional dialog perplexity baseline with a large margin on predicting Reddit feedback.
-
23 Jun 2020 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research.
-
24 Jul 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multi-reference evaluation.
-
21 Jun 2019 2 repositories listedTo investigate the strengths of this novel metric and interactive evaluation in comparison to state-of-the-art metrics and human evaluation of static conversations, we perform extended experiments with a set of models,…
-
31 Aug 2024 1 repository listedWe try to figure out how the choice of context length affects the model.
-
24 May 2023 1 repository listedThese models also suffer from posterior collapse, i.
-
12 Sep 2022 1 repository listedAutomatic evaluation of open-domain dialogs remains an unsolved problem.
-
13 Jul 2022 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedAs demonstrated by GPT-3 and T5, transformers grow in capability as parameter spaces become larger and larger.
-
22 Jun 2022 1 repository listedWe introduce GODEL (Grounded Open Dialogue Language Model), a large pre-trained language model for dialog.
-
29 May 2022 1 repository listedFinally, we provide baseline systems for these tasks and consider the function of speakers' personalities and emotions on conversation.
-
25 May 2022 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedWe introduce InstructDial, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets.
-
25 Mar 2022 1 repository listedExisting model-based metrics for system response evaluation are trained on human annotated data, which is cumbersome to collect.
-
1 Oct 2021 1 repository listedHumans often employ figurative language use in communication, including during interactions with dialog systems.
-
13 Jun 2021 1 repository listedWe instead achieve strong alignment by simultaneously modifying both the pre-trained model and the formulation of the downstream task, which is more efficient and preserves the scalability of transfer learning.
-
5 Jun 2021 1 repository listedMultiple different responses are often plausible for a given open domain dialog context.
-
1 Jun 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedOpen-domain dialog systems have a user-centric goal: to provide humans with an engaging conversation experience.
-
18 May 2021 1 repository listedHowever, existing methods for empathetic response generation usually either consider only one empathy factor or ignore the hierarchical relationships between different factors, leading to a weak ability of empathy…
-
18 Aug 2020 1 repository listedLarge end-to-end neural open-domain chatbots are becoming increasingly popular.
-
7 Jun 2020 1 repository listedThe predominant approach to open-domain dialog generation relies on end-to-end training of neural models on chat datasets.
-
21 May 2020 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedOur experiments show that CMADE achieves 89.
-
1 May 2020 1 repository listedThe lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research.
-
20 Apr 2020 1 repository listedAutomatic speech recognition (ASR) via call is essential for various applications, including AI for contact center (AICC) services.
-
17 Sep 2019 1 repository listedOpen-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or…
-
15 Aug 2019 1 repository listedOpen-domain dialog systems (also known as chatbots) have increasingly drawn attention in natural language processing.
-
1 Jul 2019 1 repository listedLarge-scale pretrained language models define state of the art in natural language processing, achieving outstanding performance on a variety of tasks.
-
30 Jun 2019 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedMost deep reinforcement learning (RL) systems are not able to learn effectively from off-policy data, especially if they cannot explore online in the environment.
-
6 Apr 2019 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections