{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modular-multitask-reinforcement-learning-with","title":"Modular Multitask Reinforcement Learning with Policy Sketches","arxiv_id":"1611.01796","date":"2016-11-06","proceeding":"ICML 2017 8","authors":["Jacob Andreas","Dan Klein","Sergey Levine"],"abstract":"We describe a framework for multitask deep reinforcement learning guided by\npolicy sketches. Sketches annotate tasks with sequences of named subtasks,\nproviding information about high-level structural relationships among tasks but\nnot how to implement them---specifically not providing the detailed guidance\nused by much previous work on learning policy abstractions for RL (e.g.\nintermediate rewards, subtask completion signals, or intrinsic motivations). To\nlearn from sketches, we present a model that associates every subtask with a\nmodular subpolicy, and jointly maximizes reward over full task-specific\npolicies by tying parameters across shared subpolicies. Optimization is\naccomplished via a decoupled actor--critic training objective that facilitates\nlearning common behaviors from multiple dissimilar reward functions. We\nevaluate the effectiveness of our approach in three environments featuring both\ndiscrete and continuous control, and with sparse rewards that can be obtained\nonly after completing a number of high-level subgoals. Experiments show that\nusing our approach to learn policies guided by sketches gives better\nperformance than existing techniques for learning task-specific or shared\npolicies, while naturally inducing a library of interpretable primitive\nbehaviors that can be recombined to rapidly adapt to new tasks.","url_abs":"http://arxiv.org/abs/1611.01796v2","url_pdf":"http://arxiv.org/pdf/1611.01796v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modular-multitask-reinforcement-learning-with","repo_url":"https://github.com/jacobandreas/psketch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"modular-multitask-reinforcement-learning-with","repo_url":"https://github.com/feryal/craft-env","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.01796","atlas_url":"https://app.syntology.ai/?focus=1611.01796","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1611.01796"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jacobandreas/psketch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/feryal/craft-env","reach":null}],"summary":{"ran_fixture":1,"ran_draft_wrong":1},"by_repo_kind":{"listed":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"b0299f082c057eee","entry":"neighbors","repo":"feryal/craft-env","repo_kind":"listed","path":"craft.py","file_url":"https://github.com/feryal/craft-env/blob/HEAD/craft.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b0299f082c057eee"}},{"code_sha256_prefix":"4e2a958dd9e63ea5","entry":"random_free","repo":"feryal/craft-env","repo_kind":"listed","path":"craft.py","file_url":"https://github.com/feryal/craft-env/blob/HEAD/craft.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4e2a958dd9e63ea5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}