{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maximum-entropy-deep-inverse-reinforcement","title":"Maximum Entropy Deep Inverse Reinforcement Learning","arxiv_id":"1507.04888","date":"2015-07-17","proceeding":null,"authors":["Markus Wulfmeier","Peter Ondruska","Ingmar Posner"],"abstract":"This paper presents a general framework for exploiting the representational\ncapacity of neural networks to approximate complex, nonlinear reward functions\nin the context of solving the inverse reinforcement learning (IRL) problem. We\nshow in this context that the Maximum Entropy paradigm for IRL lends itself\nnaturally to the efficient training of deep architectures. At test time, the\napproach leads to a computational complexity independent of the number of\ndemonstrations, which makes it especially well-suited for applications in\nlife-long learning scenarios. Our approach achieves performance commensurate to\nthe state-of-the-art on existing benchmarks while exceeding on an alternative\nbenchmark based on highly varying reward structures. Finally, we extend the\nbasic architecture - which is equivalent to a simplified subclass of Fully\nConvolutional Neural Networks (FCNNs) with width one - to include larger\nconvolutions in order to eliminate dependency on precomputed spatial features\nand work on raw input representations.","url_abs":"http://arxiv.org/abs/1507.04888v3","url_pdf":"http://arxiv.org/pdf/1507.04888v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"maximum-entropy-deep-inverse-reinforcement","repo_url":"https://github.com/XiWen0627/MaxEnIRLinCycling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"maximum-entropy-deep-inverse-reinforcement","repo_url":"https://github.com/yrlu/irl-imitation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1507.04888","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1507.04888"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yrlu/irl-imitation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/XiWen0627/MaxEnIRLinCycling","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"98c2dba45c08a2d0","entry":"load_road_network","repo":"XiWen0627/MaxEnIRLinCycling","repo_kind":"listed","path":"MapMatching/mapMatching.py","file_url":"https://github.com/XiWen0627/MaxEnIRLinCycling/blob/HEAD/MapMatching/mapMatching.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"98c2dba45c08a2d0"}},{"code_sha256_prefix":"9e2bd4fb3cec735a","entry":"optimal_value","repo":"XiWen0627/MaxEnIRLinCycling","repo_kind":"listed","path":"IRL/DPforGridBike_v1.py","file_url":"https://github.com/XiWen0627/MaxEnIRLinCycling/blob/HEAD/IRL/DPforGridBike_v1.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9e2bd4fb3cec735a"}},{"code_sha256_prefix":"78b901141fad305b","entry":"to_pixels","repo":"XiWen0627/MaxEnIRLinCycling","repo_kind":"listed","path":"MapMatching/mapMatching.py","file_url":"https://github.com/XiWen0627/MaxEnIRLinCycling/blob/HEAD/MapMatching/mapMatching.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"78b901141fad305b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}