{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gradual-transition-from-bellman-optimality","title":"Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning","arxiv_id":"2506.05968","date":"2025-06-06","proceeding":null,"authors":["Motoki Omura","Kazuki Ota","Takayuki Osa","Yusuke Mukuta","Tatsuya Harada"],"abstract":"For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q-values for the current policy using the Bellman operator. These algorithms for continuous actions rely exclusively on policy updates for improvement, which often results in low sample efficiency. This study examines the effectiveness of incorporating the Bellman optimality operator into actor-critic frameworks. Experiments in a simple environment show that modeling optimal values accelerates learning but leads to overestimation bias. To address this, we propose an annealing approach that gradually transitions from the Bellman optimality operator to the Bellman operator, thereby accelerating learning while mitigating bias. Our method, combined with TD3 and SAC, significantly outperforms existing approaches across various locomotion and manipulation tasks, demonstrating improved performance and robustness to hyperparameters related to optimality.","url_abs":"https://arxiv.org/abs/2506.05968v1","url_pdf":"https://arxiv.org/pdf/2506.05968v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gradual-transition-from-bellman-optimality","repo_url":"https://github.com/motokiomura/annealed-q-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"jax","reach":null}],"tasks":[{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"clipped-double-q-learning","method_name":"Clipped Double Q-learning"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"sac","method_name":"SAC"},{"method_slug":"td3","method_name":"TD3"},{"method_slug":"target-policy-smoothing","method_name":"Target Policy Smoothing"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.05968","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.05968"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/motokiomura/annealed-q-learning","reach":null}],"summary":{"ran_violates":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"db633b3fb7506b8f","entry":"iql_loss","repo":"motokiomura/annealed-q-learning","repo_kind":"official","path":"jaxrl/agents/sac/critic.py","file_url":"https://github.com/motokiomura/annealed-q-learning/blob/HEAD/jaxrl/agents/sac/critic.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"db633b3fb7506b8f"}},{"code_sha256_prefix":"ee22ca8c40f709c2","entry":"Model","repo":"motokiomura/annealed-q-learning","repo_kind":"official","path":"jaxrl/agents/sac/critic.py","file_url":"https://github.com/motokiomura/annealed-q-learning/blob/HEAD/jaxrl/agents/sac/critic.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ee22ca8c40f709c2"}},{"code_sha256_prefix":"084989881c394957","entry":"restore_checkpoint_if_existing","repo":"motokiomura/annealed-q-learning","repo_kind":"official","path":"train_parallel.py","file_url":"https://github.com/motokiomura/annealed-q-learning/blob/HEAD/train_parallel.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"084989881c394957"}},{"code_sha256_prefix":"37b713cf8d32aec8","entry":"update_iql","repo":"motokiomura/annealed-q-learning","repo_kind":"official","path":"jaxrl/agents/sac/critic.py","file_url":"https://github.com/motokiomura/annealed-q-learning/blob/HEAD/jaxrl/agents/sac/critic.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"37b713cf8d32aec8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}