Methods › Natural Language Processing › Dialog System Evaluation › ENIGMA

ENIGMA

7 papers tagged archive 2025-07-28

Introduced by Haoming Jiang et al. in Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

ENIGMA is an evaluation framework for dialog systems based on Pearson and Spearman's rank correlations between the estimated rewards and the true rewards. ENIGMA only requires a handful of pre-collected experience data, and therefore does not involve human interaction with the target policy during the evaluation, making automatic evaluations feasible. More importantly, ENIGMA is model-free and agnostic to the behavior policies for collecting the experience data (see details in Section 2), which significantly alleviates the technical difficulties of modeling complex dialogue environments and human behaviors.

PaperSource

Papers archive 2025-07-28

7 shown of 7, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

11 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Automated Theorem Proving2
Classification1
Diagnostic1
GPU1
Graph Neural Network1
Model-based Reinforcement Learning1
Off-policy evaluation1
Reinforcement Learning1
Reinforcement Learning (RL)1
Text Generation1
reinforcement-learning1

Usage over time archive 2025-07-28

Papers per year tagged with ENIGMA: 2021 to 2023, peak 4 4 0 2021: 4 papers 2021 2022: 2 papers 2022 2023: 1 paper 2023
Papers per year the archive tags with this method, by the paper's archive date (7 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Dialog System Evaluation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections