{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/solving-compositional-reinforcement-learning-1","title":"Solving Compositional Reinforcement Learning Problems via Task Reduction","arxiv_id":"2103.07607","date":"2021-03-13","proceeding":"ICLR 2021 1","authors":["Yunfei Li","Yilin Wu","Huazhe Xu","Xiaolong Wang","Yi Wu"],"abstract":"We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction and self-imitation. Task reduction tackles a hard-to-solve task by actively reducing it to an easier task whose solution is known by the RL agent. Once the original hard task is successfully solved by task reduction, the agent naturally obtains a self-generated solution trajectory to imitate. By continuously collecting and imitating such demonstrations, the agent is able to progressively expand the solved subspace in the entire task space. Experiment results show that SIR can significantly accelerate and improve learning on a variety of challenging sparse-reward continuous-control problems with compositional structures. Code and videos are available at https://sites.google.com/view/sir-compositional.","url_abs":"https://arxiv.org/abs/2103.07607v2","url_pdf":"https://arxiv.org/pdf/2103.07607v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"solving-compositional-reinforcement-learning-1","repo_url":"https://github.com/IrisLi17/self-imitation-via-reduction","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2103.07607","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2103.07607"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IrisLi17/self-imitation-via-reduction","reach":null}],"summary":{"ran_draft_wrong":2,"ran_violates":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"62394c3c1449402d","entry":"constfn","repo":"IrisLi17/self-imitation-via-reduction","repo_kind":"official","path":"baselines/ppo_sir/ppo_sir.py","file_url":"https://github.com/IrisLi17/self-imitation-via-reduction/blob/HEAD/baselines/ppo_sir/ppo_sir.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"62394c3c1449402d"}},{"code_sha256_prefix":"e9a0454914deaf73","entry":"get_schedule_fn","repo":"IrisLi17/self-imitation-via-reduction","repo_kind":"official","path":"baselines/ppo_sir/ppo_sir.py","file_url":"https://github.com/IrisLi17/self-imitation-via-reduction/blob/HEAD/baselines/ppo_sir/ppo_sir.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e9a0454914deaf73"}},{"code_sha256_prefix":"abd71f89b632da7f","entry":"swap_and_flatten","repo":"IrisLi17/self-imitation-via-reduction","repo_kind":"official","path":"baselines/ppo_sir/ppo_sir.py","file_url":"https://github.com/IrisLi17/self-imitation-via-reduction/blob/HEAD/baselines/ppo_sir/ppo_sir.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"abd71f89b632da7f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}