{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-robust-rewards-with-adversarial","title":"Learning Robust Rewards with Adversarial Inverse Reinforcement Learning","arxiv_id":"1710.11248","date":"2017-10-30","proceeding":null,"authors":["Justin Fu","Katie Luo","Sergey Levine"],"abstract":"Reinforcement learning provides a powerful and general framework for decision\nmaking and control, but its application in practice is often hindered by the\nneed for extensive feature and reward engineering. Deep reinforcement learning\nmethods can remove the need for explicit engineering of policy or value\nfeatures, but still require a manually specified reward function. Inverse\nreinforcement learning holds the promise of automatic reward acquisition, but\nhas proven exceptionally difficult to apply to large, high-dimensional problems\nwith unknown dynamics. In this work, we propose adverserial inverse\nreinforcement learning (AIRL), a practical and scalable inverse reinforcement\nlearning algorithm based on an adversarial reward learning formulation. We\ndemonstrate that AIRL is able to recover reward functions that are robust to\nchanges in dynamics, enabling us to learn policies even under significant\nvariation in the environment seen during training. Our experiments show that\nAIRL greatly outperforms prior methods in these transfer settings.","url_abs":"http://arxiv.org/abs/1710.11248v2","url_pdf":"http://arxiv.org/pdf/1710.11248v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/Div99/IQ-Learn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/Kaixhin/imitation-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/evieq01/oodil","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/ku2482/gail-airl-ppo.pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/rohitrango/Reward-bias-in-GAIL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/twni2016/f-IRL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-robust-rewards-with-adversarial","repo_url":"https://github.com/mugoh/rl-base/tree/master/rlbase/aiRL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/mujoco-games-on-ant","task":"MuJoCo Games","dataset":"Ant","model":"AIRL Fu et al. (2017)","rank_in_archive_order":3,"of":3,"metrics":{"Average Return":"127.61"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.11248","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1710.11248"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Div99/IQ-Learn","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Kaixhin/imitation-learning","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rohitrango/Reward-bias-in-GAIL","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/twni2016/f-IRL","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/evieq01/oodil","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mugoh/rl-base/tree/master/rlbase/aiRL","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ku2482/gail-airl-ppo.pytorch","reach":null}],"summary":{"unverified":5},"by_repo_kind":{"listed":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"4f15d01c0336d4d5","entry":"check_last_name","repo":"rohitrango/Reward-bias-in-GAIL","repo_kind":"listed","path":"plot_graphs.py","file_url":"https://github.com/rohitrango/Reward-bias-in-GAIL/blob/HEAD/plot_graphs.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"4f15d01c0336d4d5"}},{"code_sha256_prefix":"41fbf53df6ae4012","entry":"linear_beta_schedule","repo":"rohitrango/Reward-bias-in-GAIL","repo_kind":"listed","path":"src/imitation/algorithms/dagger.py","file_url":"https://github.com/rohitrango/Reward-bias-in-GAIL/blob/HEAD/src/imitation/algorithms/dagger.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"41fbf53df6ae4012"}},{"code_sha256_prefix":"eaea6e63c7b89d35","entry":"mce_irl","repo":"rohitrango/Reward-bias-in-GAIL","repo_kind":"listed","path":"src/imitation/algorithms/tabular_irl.py","file_url":"https://github.com/rohitrango/Reward-bias-in-GAIL/blob/HEAD/src/imitation/algorithms/tabular_irl.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"eaea6e63c7b89d35"}},{"code_sha256_prefix":"5d423a53ae972b3e","entry":"mce_occupancy_measures","repo":"rohitrango/Reward-bias-in-GAIL","repo_kind":"listed","path":"src/imitation/algorithms/tabular_irl.py","file_url":"https://github.com/rohitrango/Reward-bias-in-GAIL/blob/HEAD/src/imitation/algorithms/tabular_irl.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"5d423a53ae972b3e"}},{"code_sha256_prefix":"91b61186c976f6ac","entry":"mce_partition_fh","repo":"rohitrango/Reward-bias-in-GAIL","repo_kind":"listed","path":"src/imitation/algorithms/tabular_irl.py","file_url":"https://github.com/rohitrango/Reward-bias-in-GAIL/blob/HEAD/src/imitation/algorithms/tabular_irl.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"91b61186c976f6ac"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}