{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/computational-benefits-of-intermediate","title":"Computational Benefits of Intermediate Rewards for Goal-Reaching Policy Learning","arxiv_id":"2107.03961","date":"2021-07-08","proceeding":null,"authors":["Yuexiang Zhai","Christina Baek","Zhengyuan Zhou","Jiantao Jiao","Yi Ma"],"abstract":"Many goal-reaching reinforcement learning (RL) tasks have empirically verified that rewarding the agent on subgoals improves convergence speed and practical performance. We attempt to provide a theoretical framework to quantify the computational benefits of rewarding the completion of subgoals, in terms of the number of synchronous value iterations. In particular, we consider subgoals as one-way {\\em intermediate states}, which can only be visited once per episode and propose two settings that consider these one-way intermediate states: the one-way single-path (OWSP) and the one-way multi-path (OWMP) settings. In both OWSP and OWMP settings, we demonstrate that adding {\\em intermediate rewards} to subgoals is more computationally efficient than only rewarding the agent once it completes the goal of reaching a terminal state. We also reveal a trade-off between computational complexity and the pursuit of the shortest path in the OWMP setting: adding intermediate rewards significantly reduces the computational complexity of reaching the goal but the agent may not find the shortest path, whereas with sparse terminal rewards, the agent finds the shortest path at a significantly higher computational cost. We also corroborate our theoretical results with extensive experiments on the MiniGrid environments using Q-learning and some popular deep RL algorithms.","url_abs":"https://arxiv.org/abs/2107.03961v5","url_pdf":"https://arxiv.org/pdf/2107.03961v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"computational-benefits-of-intermediate","repo_url":"https://github.com/kebaek/minigrid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"hierarchical-reinforcement-learning","task_name":"Hierarchical Reinforcement Learning"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2107.03961","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2107.03961"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kebaek/minigrid","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"beb7d468d5185af5","entry":"reject_next_to","repo":"kebaek/minigrid","repo_kind":"official","path":"gym-minigrid/gym_minigrid/roomgrid.py","file_url":"https://github.com/kebaek/minigrid/blob/HEAD/gym-minigrid/gym_minigrid/roomgrid.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"beb7d468d5185af5"}},{"code_sha256_prefix":"59df34732c3e3418","entry":"downsample","repo":"kebaek/minigrid","repo_kind":"official","path":"gym-minigrid/gym_minigrid/rendering.py","file_url":"https://github.com/kebaek/minigrid/blob/HEAD/gym-minigrid/gym_minigrid/rendering.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"59df34732c3e3418"}},{"code_sha256_prefix":"bd9c272b32f75522","entry":"fill_coords","repo":"kebaek/minigrid","repo_kind":"official","path":"gym-minigrid/gym_minigrid/rendering.py","file_url":"https://github.com/kebaek/minigrid/blob/HEAD/gym-minigrid/gym_minigrid/rendering.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bd9c272b32f75522"}},{"code_sha256_prefix":"8990d1cc61cec299","entry":"rotate_fn","repo":"kebaek/minigrid","repo_kind":"official","path":"gym-minigrid/gym_minigrid/rendering.py","file_url":"https://github.com/kebaek/minigrid/blob/HEAD/gym-minigrid/gym_minigrid/rendering.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8990d1cc61cec299"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}