{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-trust-bellman-updates-selective","title":"Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL","arxiv_id":"2505.19923","date":"2025-05-26","proceeding":null,"authors":["Qin-Wen Luo","Ming-Kun Xie","Ye-Wen Wang","Sheng-Jun Huang"],"abstract":"Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset. To alleviate extrapolation errors, existing studies often uniformly regularize the value function or policy updates across all states. However, due to substantial variations in data quality, the fixed regularization strength often leads to a dilemma: Weak regularization strength fails to address extrapolation errors and value overestimation, while strong regularization strength shifts policy learning toward behavior cloning, impeding potential performance enabled by Bellman updates. To address this issue, we propose the selective state-adaptive regularization method for offline RL. Specifically, we introduce state-adaptive regularization coefficients to trust state-level Bellman-driven results, while selectively applying regularization on high-quality actions, aiming to avoid performance degradation caused by tight constraints on low-quality actions. By establishing a connection between the representative value regularization method, CQL, and explicit policy constraint methods, we effectively extend selective state-adaptive regularization to these two mainstream offline RL approaches. Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art approaches in both offline and offline-to-online settings on the D4RL benchmark.","url_abs":"https://arxiv.org/abs/2505.19923v1","url_pdf":"https://arxiv.org/pdf/2505.19923v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-trust-bellman-updates-selective","repo_url":"https://github.com/qinwenluo/ssar","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"d4rl","task_name":"D4RL"},{"task_slug":"offline-rl","task_name":"Offline RL"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2505.19923","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.19923"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/QinwenLuo/SSAR","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qinwenluo/ssar","reach":{"status":"ok"}}],"summary":{"ran":2,"ran_draft_wrong":2,"ran_violates":2,"unverified":1},"by_repo_kind":{"official":{"samples":7,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"66f4d6b9de20db19","entry":"ParaNet","repo":"qinwenluo/ssar","repo_kind":"official","path":"cql.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/cql.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"66f4d6b9de20db19"}},{"code_sha256_prefix":"d7ebf290cb024d50","entry":"Scalar","repo":"qinwenluo/ssar","repo_kind":"official","path":"cql.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/cql.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d7ebf290cb024d50"}},{"code_sha256_prefix":"2c8a1a60cc1d1a9e","entry":"load_train_config_auto","repo":"qinwenluo/ssar","repo_kind":"official","path":"cql.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/cql.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2c8a1a60cc1d1a9e"}},{"code_sha256_prefix":"8bcabffc920c479a","entry":"load_train_config_auto","repo":"qinwenluo/ssar","repo_kind":"official","path":"td3_bc.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/td3_bc.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8bcabffc920c479a"}},{"code_sha256_prefix":"4a7997791dd920cf","entry":"modify_reward_online","repo":"qinwenluo/ssar","repo_kind":"official","path":"cql.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/cql.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4a7997791dd920cf"}},{"code_sha256_prefix":"88fb609b188e3ee3","entry":"modify_reward_online","repo":"qinwenluo/ssar","repo_kind":"official","path":"td3_bc.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/td3_bc.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"88fb609b188e3ee3"}},{"code_sha256_prefix":"edaba7b96a876979","entry":"ContinuousCQL","repo":"qinwenluo/ssar","repo_kind":"official","path":"cql.py","file_url":"https://github.com/qinwenluo/ssar/blob/HEAD/cql.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"edaba7b96a876979"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}