{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/re-evaluating-evaluation","title":"Re-evaluating Evaluation","arxiv_id":"1806.02643","date":"2018-06-07","proceeding":"NeurIPS 2018 12","authors":["David Balduzzi","Karl Tuyls","Julien Perolat","Thore Graepel"],"abstract":"Progress in machine learning is measured by careful evaluation on problems of\noutstanding common interest. However, the proliferation of benchmark suites and\nenvironments, adversarial attacks, and other complications has diluted the\nbasic evaluation model by overwhelming researchers with choices. Deliberate or\naccidental cherry picking is increasingly likely, and designing well-balanced\nevaluation suites requires increasing effort. In this paper we take a step back\nand propose Nash averaging. The approach builds on a detailed analysis of the\nalgebraic structure of evaluation in two basic scenarios: agent-vs-agent and\nagent-vs-task. The key strength of Nash averaging is that it automatically\nadapts to redundancies in evaluation data, so that results are not biased by\nthe incorporation of easy tasks or weak agents. Nash averaging thus encourages\nmaximally inclusive evaluation -- since there is no harm (computational cost\naside) from including all available tasks and agents.","url_abs":"http://arxiv.org/abs/1806.02643v2","url_pdf":"http://arxiv.org/pdf/1806.02643v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"re-evaluating-evaluation","repo_url":"https://github.com/PhilipFelizarta/Maxent-Nash-Implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"re-evaluating-evaluation","repo_url":"https://github.com/deepmind/open_spiel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.02643","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}