{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/refining-minimax-regret-for-unsupervised","title":"Refining Minimax Regret for Unsupervised Environment Design","arxiv_id":"2402.12284","date":"2024-02-19","proceeding":null,"authors":["Michael Beukman","Samuel Coward","Michael Matthews","Mattie Fellows","Minqi Jiang","Michael Dennis","Jakob Foerster"],"abstract":"In unsupervised environment design, reinforcement learning agents are trained on environment configurations (levels) generated by an adversary that maximises some objective. Regret is a commonly used objective that theoretically results in a minimax regret (MMR) policy with desirable robustness guarantees; in particular, the agent's maximum regret is bounded. However, once the agent reaches this regret bound on all levels, the adversary will only sample levels where regret cannot be further reduced. Although there are possible performance improvements to be made outside of these regret-maximising levels, learning stagnates. In this work, we introduce Bayesian level-perfect MMR (BLP), a refinement of the minimax regret objective that overcomes this limitation. We formally show that solving for this objective results in a subset of MMR policies, and that BLP policies act consistently with a Perfect Bayesian policy over all levels. We further introduce an algorithm, ReMiDi, that results in a BLP policy at convergence. We empirically demonstrate that training on levels from a minimax regret adversary causes learning to prematurely stagnate, but that ReMiDi continues learning.","url_abs":"https://arxiv.org/abs/2402.12284v2","url_pdf":"https://arxiv.org/pdf/2402.12284v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"refining-minimax-regret-for-unsupervised","repo_url":"https://github.com/michael-beukman/remidi","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.12284","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.12284"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/michael-beukman/remidi","reach":{"status":"ok"}}],"summary":{"ran":6,"unverified":2},"by_repo_kind":{"official":{"samples":8,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"f20aa0108a055874","entry":"compute_gae","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/common/ppo.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/common/ppo.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f20aa0108a055874"}},{"code_sha256_prefix":"807cd87fb87d6b5d","entry":"fill_coords","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/environments/maze/vis.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/environments/maze/vis.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"807cd87fb87d6b5d"}},{"code_sha256_prefix":"a0fcd1a9937f59d6","entry":"load_compressed_pickle","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/utils.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a0fcd1a9937f59d6"}},{"code_sha256_prefix":"2738cb883a4907b2","entry":"make_lever_level_generator","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/environments/lever_game/env.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/environments/lever_game/env.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2738cb883a4907b2"}},{"code_sha256_prefix":"27e5ee93023a1020","entry":"point_in_rect","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/environments/maze/vis.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/environments/maze/vis.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"27e5ee93023a1020"}},{"code_sha256_prefix":"680139aaead3ac7d","entry":"point_in_triangle","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/environments/maze/vis.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/environments/maze/vis.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"680139aaead3ac7d"}},{"code_sha256_prefix":"b51283140c732791","entry":"sample_trajectories_rnn","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/common/ppo.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/common/ppo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b51283140c732791"}},{"code_sha256_prefix":"0b73880de813328a","entry":"update_actor_critic_rnn","repo":"michael-beukman/remidi","repo_kind":"official","path":"lib/common/ppo.py","file_url":"https://github.com/michael-beukman/remidi/blob/HEAD/lib/common/ppo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0b73880de813328a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}