{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/state-advantage-weighting-for-offline-rl","title":"State Advantage Weighting for Offline RL","arxiv_id":"2210.04251","date":"2022-10-09","proceeding":null,"authors":["Jiafei Lyu","Aicheng Gong","Le Wan","Zongqing Lu","Xiu Li"],"abstract":"We present state advantage weighting for offline reinforcement learning (RL). In contrast to action advantage $A(s,a)$ that we commonly adopt in QSA learning, we leverage state advantage $A(s,s^\\prime)$ and QSS learning for offline RL, hence decoupling the action from values. We expect the agent can get to the high-reward state and the action is determined by how the agent can get to that corresponding state. Experiments on D4RL datasets show that our proposed method can achieve remarkable performance against the common baselines. Furthermore, our method shows good generalization capability when transferring from offline to online.","url_abs":"https://arxiv.org/abs/2210.04251v2","url_pdf":"https://arxiv.org/pdf/2210.04251v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"d4rl","task_name":"D4RL"},{"task_slug":"offline-rl","task_name":"Offline RL"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2210.04251","atlas_url":"https://app.syntology.ai/?focus=2210.04251","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.04251"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/aviralkumar2907/CQL","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/apple/ml-uwac","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/ikostrikov/implicit_q_learning","reach":null}],"summary":{"unverified":5},"by_repo_kind":{"found_in_text":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"063042eb56bcac07","entry":"Model","repo":"ikostrikov/implicit_q_learning","repo_kind":"found_in_text","path":"learner.py","file_url":"https://github.com/ikostrikov/implicit_q_learning/blob/HEAD/learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"063042eb56bcac07"}},{"code_sha256_prefix":"f3378f6aee6adf90","entry":"_update_jit","repo":"ikostrikov/implicit_q_learning","repo_kind":"found_in_text","path":"learner.py","file_url":"https://github.com/ikostrikov/implicit_q_learning/blob/HEAD/learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f3378f6aee6adf90"}},{"code_sha256_prefix":"87e5a6628edb14db","entry":"target_update","repo":"ikostrikov/implicit_q_learning","repo_kind":"found_in_text","path":"learner.py","file_url":"https://github.com/ikostrikov/implicit_q_learning/blob/HEAD/learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"87e5a6628edb14db"}},{"code_sha256_prefix":"8667d42b5ba48077","entry":"update_q","repo":"ikostrikov/implicit_q_learning","repo_kind":"found_in_text","path":"learner.py","file_url":"https://github.com/ikostrikov/implicit_q_learning/blob/HEAD/learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8667d42b5ba48077"}},{"code_sha256_prefix":"76207004807550fe","entry":"update_v","repo":"ikostrikov/implicit_q_learning","repo_kind":"found_in_text","path":"learner.py","file_url":"https://github.com/ikostrikov/implicit_q_learning/blob/HEAD/learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"76207004807550fe"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}