Methods › Natural Language Processing › Dialog System Evaluation › ENIGMA
ENIGMA
Introduced by Haoming Jiang et al. in Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
ENIGMA is an evaluation framework for dialog systems based on Pearson and Spearman's rank correlations between the estimated rewards and the true rewards. ENIGMA only requires a handful of pre-collected experience data, and therefore does not involve human interaction with the target policy during the evaluation, making automatic evaluations feasible. More importantly, ENIGMA is model-free and agnostic to the behavior policies for collecting the experience data (see details in Section 2), which significantly alleviates the technical difficulties of modeling complex dialogue environments and human behaviors.
Papers archive 2025-07-28
7 shown of 7, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
MizAR 60 for Mizar 50 12 Mar 2023 · 0 repositories · arXiv:2303.06686
-
Multi-site benchmark classification of major depressive disorder using machine learning on cortical and subcortical measures 16 Jun 2022 · 0 repositories · arXiv:2206.08122
-
The Isabelle ENIGMA 4 May 2022 · 1 repository · arXiv:2205.01981
-
Learning Theorem Proving Components 21 Jul 2021 · 1 repository · arXiv:2107.10034
-
Fast and Slow Enigmas and Parental Guidance 14 Jul 2021 · 1 repository · arXiv:2107.06750
-
Improving ENIGMA-Style Clause Selection While Learning From History 26 Feb 2021 · 0 repositories · arXiv:2102.13564
-
Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach 20 Feb 2021 · 1 repository · arXiv:2102.10242
Tasks archive 2025-07-28
11 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections