{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evil-evolution-strategies-for-generalisable","title":"EvIL: Evolution Strategies for Generalisable Imitation Learning","arxiv_id":"2406.11905","date":"2024-06-15","proceeding":null,"authors":["Silvia Sapora","Gokul Swamy","Chris Lu","Yee Whye Teh","Jakob Nicolaus Foerster"],"abstract":"Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g. demonstrations collected in simulation but deployment in the real world). Compared to policy-centric approaches to IL like behavioural cloning, reward-centric approaches like inverse reinforcement learning (IRL) often better replicate expert behaviour in new environments. This transfer is usually performed by optimising the recovered reward under the dynamics of the target environment. However, (a) we find that modern deep IL algorithms frequently recover rewards which induce policies far weaker than the expert, even in the same environment the demonstrations were collected in. Furthermore, (b) these rewards are often quite poorly shaped, necessitating extensive environment interaction to optimise effectively. We provide simple and scalable fixes to both of these concerns. For (a), we find that reward model ensembles combined with a slightly different training objective significantly improves re-training and transfer performance. For (b), we propose a novel evolution-strategies based method EvIL to optimise for a reward-shaping term that speeds up re-training in the target environment, closing a gap left open by the classical theory of IRL. On a suite of continuous control tasks, we are able to re-train policies in target (and source) environments more interaction-efficiently than prior work.","url_abs":"https://arxiv.org/abs/2406.11905v1","url_pdf":"https://arxiv.org/pdf/2406.11905v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evil-evolution-strategies-for-generalisable","repo_url":"https://github.com/SilviaSapora/evil","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"jax","reach":{"status":"ok"}}],"tasks":[{"task_slug":"behavioural-cloning","task_name":"Behavioural cloning"},{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.11905","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.11905"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/SilviaSapora/evil","reach":{"status":"ok"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"af6757bd063d8a4c","entry":"get_new_found_array","repo":"SilviaSapora/evil","repo_kind":"official","path":"evil/envs/gridworld.py","file_url":"https://github.com/SilviaSapora/evil/blob/HEAD/evil/envs/gridworld.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"af6757bd063d8a4c"}},{"code_sha256_prefix":"5afb2039182409d3","entry":"get_state_dict_from_obs","repo":"SilviaSapora/evil","repo_kind":"official","path":"evil/envs/gridworld.py","file_url":"https://github.com/SilviaSapora/evil/blob/HEAD/evil/envs/gridworld.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5afb2039182409d3"}},{"code_sha256_prefix":"7695e218fc5bced0","entry":"get_state_from_obs","repo":"SilviaSapora/evil","repo_kind":"official","path":"evil/envs/gridworld.py","file_url":"https://github.com/SilviaSapora/evil/blob/HEAD/evil/envs/gridworld.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7695e218fc5bced0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}