Papers › USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

1 May 2020ACL 2020 6arXiv:2005.00456archive 2025-07-28

Shikib Mehri, Maxine Eskenazi

The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog. USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog. USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, system-level: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0). USR additionally produces interpretable measures for several desirable properties of dialog.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

shikib/usr officialmentioned in paperpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Dialogue EvaluationOpen-Domain DialogText Generation

Datasets

Introduced by this paper, per the archive.

USR-PersonaChatUSR-TopicalChat

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Dialogue Evaluation USR-PersonaChat USR - DR (x = c) Pearson Correlation 0.6087 #2 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR - DR (x = c) Spearman Correlation 0.4814 #2 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR Pearson Correlation 0.4115 #3 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR Spearman Correlation 0.4693 #3 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR - MLM Pearson Correlation 0.0788 #4 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR - MLM Spearman Correlation 0.0795 #4 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR - DR (x = f) Pearson Correlation -0.0454 #5 of 5 Archive leaderboard report
Dialogue Evaluation USR-PersonaChat USR - DR (x = f) Spearman Correlation -0.0495 #5 of 5 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR Pearson Correlation 0.4220 #3 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR Spearman Correlation 0.4192 #3 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - DR (x = c) Pearson Correlation 0.4068 #4 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - DR (x = c) Spearman Correlation 0.3245 #4 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - MLM Pearson Correlation 0.3345 #5 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - MLM Spearman Correlation 0.3086 #5 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - DR (x = f) Pearson Correlation 0.3221 #6 of 6 Archive leaderboard report
Dialogue Evaluation USR-TopicalChat USR - DR (x = f) Spearman Correlation 0.1419 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections