{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-from-sparse-offline-datasets-via","title":"Learning from Sparse Offline Datasets via Conservative Density Estimation","arxiv_id":"2401.08819","date":"2024-01-16","proceeding":null,"authors":["Zhepeng Cen","Zuxin Liu","Zitong Wang","Yihang Yao","Henry Lam","Ding Zhao"],"abstract":"Offline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL.","url_abs":"https://arxiv.org/abs/2401.08819v2","url_pdf":"https://arxiv.org/pdf/2401.08819v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-from-sparse-offline-datasets-via","repo_url":"https://github.com/czp16/cde-offline-rl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"d4rl","task_name":"D4RL"},{"task_slug":"density-estimation","task_name":"Density Estimation"},{"task_slug":"offline-rl","task_name":"Offline RL"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2401.08819","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.08819"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/czp16/cde-offline-rl","reach":null}],"summary":{"ran":1,"ran_draft_wrong":1,"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1df7301b0c11658e","entry":"CDELearner","repo":"czp16/cde-offline-rl","repo_kind":"official","path":"cde.py","file_url":"https://github.com/czp16/cde-offline-rl/blob/HEAD/cde.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1df7301b0c11658e"}},{"code_sha256_prefix":"e06c8c099e9fd7f4","entry":"get_f_div_fn","repo":"czp16/cde-offline-rl","repo_kind":"official","path":"cde.py","file_url":"https://github.com/czp16/cde-offline-rl/blob/HEAD/cde.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e06c8c099e9fd7f4"}},{"code_sha256_prefix":"cb62199e8e3037a1","entry":"to_tensor","repo":"czp16/cde-offline-rl","repo_kind":"official","path":"cde.py","file_url":"https://github.com/czp16/cde-offline-rl/blob/HEAD/cde.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cb62199e8e3037a1"}},{"code_sha256_prefix":"731a435187750caf","entry":"Config","repo":"czp16/cde-offline-rl","repo_kind":"official","path":"cde.py","file_url":"https://github.com/czp16/cde-offline-rl/blob/HEAD/cde.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"731a435187750caf"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}