Browse State-of-the-Art › Dialogue Evaluation
Dialogue Evaluation
61 papers with code · 2 benchmarks · 8 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| USR-TopicalChat (6 rows) | MDD-Eval | MDD-Eval: Self-Training on Augmented Data for Multi-Domain... | code | — | Compare |
| USR-PersonaChat (5 rows) | Lin-Reg (all) | Proxy Indicators for the Quality of Open-domain Dialogues | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 61 papers with code (97 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jan 2017 8 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)In this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from…
-
18 Dec 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Our method is used to evaluate four state-of-the-art open-domain dialogue systems and compared with existing approaches.
-
25 Oct 2022 2 repositories listedRecent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment.
-
3 Nov 2021 2 repositories listedThe development of Open-Domain Dialogue Systems (ODS)is a trending topic due to the large number of research challenges, large societal and business impact, and advances in the underlying technology.
-
23 Jun 2020 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research.
-
4 Nov 2019 2 repositories listedIn this paper, we investigate the possibility and efficacy of estimating utterance-level engagement and define a novel metric, {\em predictive engagement}, for automatic evaluation of open-domain dialogue systems.
-
24 Jul 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multi-reference evaluation.
-
21 Jun 2019 2 repositories listedTo investigate the strengths of this novel metric and interactive evaluation in comparison to state-of-the-art metrics and human evaluation of static conversations, we perform extended experiments with a set of models,…
-
28 May 2025 1 repository listedAs the capabilities of chatbots and their underlying LLMs continue to dramatically improve, evaluating their performance has increasingly become a major blocker to their further development.
-
22 Apr 2025 1 repository listedIn this paper, we describe our participation in the RuTermEval competition devoted to extracting nested terms.
-
9 Apr 2025 1 repository listedThe participants experimented mainly with large language models in zero-shot, few-shot and fine-tuning formats.
-
7 Apr 2025 1 repository listedRecent advancements in large language models (LLMs) have revolutionized their ability to handle single-turn tasks, yet real-world applications demand sophisticated multi-turn interactions.
-
17 Jan 2025 1 repository listedOne such auxiliary loss function is Bag-of-Words (BoW) loss, defined as the cross-entropy loss for predicting all the words/tokens of the next utterance.
-
12 Jan 2025 1 repository listedAdvancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses.
-
20 Aug 2024 1 repository listedAlthough human evaluation remains the gold standard for open-domain dialogue evaluation, the growing popularity of automated evaluation using Large Language Models (LLMs) has also extended to dialogue.
-
16 Jul 2024 1 repository listedMotivated by the need for lightweight, open source, and multilingual dialogue evaluators, this paper introduces GenResCoh (Generated Responses targeting Coherence).
-
24 May 2024 1 repository listedOur approach introduces several techniques: (1) Contrastive learning to differentiate between robust and non-robust response embeddings; (2) A novel metric for semantic sensitivity that combines embedding cosine…
-
1 Apr 2024 1 repository listedTrainable evaluation metrics, typically trained with true positive and randomly selected negative responses, tend to assign higher scores to responses that share greater content similarity with a given context.
-
1 Apr 2024 1 repository listedRecent studies proposed evaluation metrics that assess generated responses by considering their relevance to previous dialogue histories.
-
24 Dec 2023 1 repository listedYet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc.
-
3 Nov 2023 1 repository listedIn this paper, we propose DialogBench, a dialogue evaluation benchmark that contains 12 dialogue tasks to probe the capabilities of LLMs as human-like dialogue systems should have.
-
13 Oct 2023 1 repository listedThe English dialogue data are extended to nine other languages with commercial machine translation systems.
-
14 Sep 2023 1 repository listedHuman evaluation has been widely accepted as the standard for evaluating chat-oriented dialogue systems.
-
31 Aug 2023 1 repository listedDespite significant research effort in the development of automatic dialogue evaluation metrics, little thought is given to evaluating dialogues other than in English.
-
31 Aug 2023 1 repository listedThe main limiting factor in the development of robust multilingual dialogue evaluation metrics is the lack of multilingual data and the limited availability of open sourced multilingual dialogue systems.
-
27 Jun 2023 1 repository listedExisting reference-free turn-level evaluation metrics for chatbots inadequately capture the interaction between the user and the system.
-
8 May 2023 1 repository listedDespite the recent advances in open-domain dialogue systems, building a reliable evaluation metric is still a challenging problem.
-
28 Feb 2023 1 repository listedWe present GLM-Dialog, a large-scale language model (LLM) with 10B parameters capable of knowledge-grounded conversation in Chinese using a search engine to access the Internet knowledge.
-
17 Aug 2022 1 repository listedThis paper introduces a novel Self-supervised Fine-grained Dialogue Evaluation framework (SelF-Eval).
-
3 Jun 2022 1 repository listedThe first task is framed as a binary classification problem.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections