{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-a-diffusion-model-policy-from","title":"Learning a Diffusion Model Policy from Rewards via Q-Score Matching","arxiv_id":"2312.11752","date":"2023-12-18","proceeding":null,"authors":["Michael Psenka","Alejandro Escontrela","Pieter Abbeel","Yi Ma"],"abstract":"Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning. This is due to their natural ability to optimize an expressive class of distributions over a continuous space. However, previous works fail to exploit the score-based structure of diffusion models, and instead utilize a simple behavior cloning term to train the actor, limiting their ability in the actor-critic setting. In this paper, we present a theoretical framework linking the structure of diffusion model policies to a learned Q-function, by linking the structure between the score of the policy to the action gradient of the Q-function. We focus on off-policy reinforcement learning and propose a new policy update method from this theory, which we denote Q-score matching. Notably, this algorithm only needs to differentiate through the denoising model rather than the entire diffusion model evaluation, and converged policies through Q-score matching are implicitly multi-modal and explorative in continuous domains. We conduct experiments in simulated environments to demonstrate the viability of our proposed method and compare to popular baselines. Source code is available from the project website: https://michaelpsenka.io/qsm.","url_abs":"https://arxiv.org/abs/2312.11752v4","url_pdf":"https://arxiv.org/pdf/2312.11752v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-a-diffusion-model-policy-from","repo_url":"https://github.com/Alescontrela/score_matching_rl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2312.11752","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.11752"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Alescontrela/score_matching_rl","reach":null}],"summary":{"ran_honours":2,"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"87d1bc442d93916d","entry":"cosine_beta_schedule","repo":"Alescontrela/score_matching_rl","repo_kind":"official","path":"jaxrl5/agents/score_matching/score_matching_learner.py","file_url":"https://github.com/Alescontrela/score_matching_rl/blob/HEAD/jaxrl5/agents/score_matching/score_matching_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"87d1bc442d93916d"}},{"code_sha256_prefix":"70750a8a4843e1c7","entry":"ddpm_sampler","repo":"Alescontrela/score_matching_rl","repo_kind":"official","path":"jaxrl5/agents/score_matching/score_matching_learner.py","file_url":"https://github.com/Alescontrela/score_matching_rl/blob/HEAD/jaxrl5/agents/score_matching/score_matching_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"70750a8a4843e1c7"}},{"code_sha256_prefix":"539aaf559a17a19a","entry":"vp_beta_schedule","repo":"Alescontrela/score_matching_rl","repo_kind":"official","path":"jaxrl5/agents/score_matching/score_matching_learner.py","file_url":"https://github.com/Alescontrela/score_matching_rl/blob/HEAD/jaxrl5/agents/score_matching/score_matching_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"539aaf559a17a19a"}},{"code_sha256_prefix":"4e662c8ce5053be2","entry":"ScoreMatchingLearner","repo":"Alescontrela/score_matching_rl","repo_kind":"official","path":"jaxrl5/agents/score_matching/score_matching_learner.py","file_url":"https://github.com/Alescontrela/score_matching_rl/blob/HEAD/jaxrl5/agents/score_matching/score_matching_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4e662c8ce5053be2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}