{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maximum-entropy-reinforcement-learning-with-1","title":"Maximum Entropy Reinforcement Learning with Diffusion Policy","arxiv_id":"2502.11612","date":"2025-02-17","proceeding":null,"authors":["Xiaoyi Dong","Jian Cheng","Xi Sheryl Zhang"],"abstract":"The Soft Actor-Critic (SAC) algorithm with a Gaussian policy has become a mainstream implementation for realizing the Maximum Entropy Reinforcement Learning (MaxEnt RL) objective, which incorporates entropy maximization to encourage exploration and enhance policy robustness. While the Gaussian policy performs well on simpler tasks, its exploration capacity and potential performance in complex multi-goal RL environments are limited by its inherent unimodality. In this paper, we employ the diffusion model, a powerful generative model capable of capturing complex multimodal distributions, as the policy representation to fulfill the MaxEnt RL objective, developing a method named MaxEnt RL with Diffusion Policy (MaxEntDP). Our method enables efficient exploration and brings the policy closer to the optimal MaxEnt policy. Experimental results on Mujoco benchmarks show that MaxEntDP outperforms the Gaussian policy and other generative models within the MaxEnt RL framework, and performs comparably to other state-of-the-art diffusion-based online RL algorithms. Our code is available at https://github.com/diffusionyes/MaxEntDP.","url_abs":"https://arxiv.org/abs/2502.11612v1","url_pdf":"https://arxiv.org/pdf/2502.11612v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"maximum-entropy-reinforcement-learning-with-1","repo_url":"https://github.com/diffusionyes/maxentdp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"jax","reach":null}],"tasks":[{"task_slug":"efficient-exploration","task_name":"Efficient Exploration"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2502.11612","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2502.11612"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/diffusionyes/MaxEntDP","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/diffusionyes/maxentdp","reach":null}],"summary":{"ran":6,"ran_fixture":2,"ran_honours":2,"unverified":3},"by_repo_kind":{"official":{"samples":13,"ran":10,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":13,"samples":[{"code_sha256_prefix":"1d2b364aa107cd48","entry":"Agent","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1d2b364aa107cd48"}},{"code_sha256_prefix":"0533319139439dda","entry":"DDPM","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0533319139439dda"}},{"code_sha256_prefix":"137f7334fa321901","entry":"FourierFeatures","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"137f7334fa321901"}},{"code_sha256_prefix":"5cb27561835800a5","entry":"MLP","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5cb27561835800a5"}},{"code_sha256_prefix":"7467a4afe03535e5","entry":"NoiseScheduleVP","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7467a4afe03535e5"}},{"code_sha256_prefix":"ded9e20f53391a22","entry":"StateActionValue","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ded9e20f53391a22"}},{"code_sha256_prefix":"d1899127f3f7bcdc","entry":"_sample_actions","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d1899127f3f7bcdc"}},{"code_sha256_prefix":"8d277cc41d953efc","entry":"interpolate_fn","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8d277cc41d953efc"}},{"code_sha256_prefix":"7fc1d127bf81537f","entry":"mish","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7fc1d127bf81537f"}},{"code_sha256_prefix":"be4c7efe187d2c0e","entry":"tensorstats","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"be4c7efe187d2c0e"}},{"code_sha256_prefix":"ce659c306357023b","entry":"MaxEntropyLearner","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ce659c306357023b"}},{"code_sha256_prefix":"e9fbe7502095bfe0","entry":"_eval_actions","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e9fbe7502095bfe0"}},{"code_sha256_prefix":"2cef9ed98e841b7f","entry":"ddpm_sampler","repo":"diffusionyes/maxentdp","repo_kind":"official","path":"jaxrl5/agents/score_matching/max_entropy_learner.py","file_url":"https://github.com/diffusionyes/maxentdp/blob/HEAD/jaxrl5/agents/score_matching/max_entropy_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2cef9ed98e841b7f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}