{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/q-pensieve-boosting-sample-efficiency-of","title":"Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots","arxiv_id":"2212.03117","date":"2022-12-06","proceeding":null,"authors":["Wei Hung","Bo-Kai Huang","Ping-Chun Hsieh","Xi Liu"],"abstract":"Many real-world continuous control problems are in the dilemma of weighing the pros and cons, multi-objective reinforcement learning (MORL) serves as a generic framework of learning control policies for different preferences over objectives. However, the existing MORL methods either rely on multiple passes of explicit search for finding the Pareto front and therefore are not sample-efficient, or utilizes a shared policy network for coarse knowledge sharing among policies. To boost the sample efficiency of MORL, we propose Q-Pensieve, a policy improvement scheme that stores a collection of Q-snapshots to jointly determine the policy update direction and thereby enables data sharing at the policy level. We show that Q-Pensieve can be naturally integrated with soft policy iteration with convergence guarantee. To substantiate this concept, we propose the technique of Q replay buffer, which stores the learned Q-networks from the past iterations, and arrive at a practical actor-critic implementation. Through extensive experiments and an ablation study, we demonstrate that with much fewer samples, the proposed algorithm can outperform the benchmark MORL methods on a variety of MORL benchmark tasks.","url_abs":"https://arxiv.org/abs/2212.03117v2","url_pdf":"https://arxiv.org/pdf/2212.03117v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"q-pensieve-boosting-sample-efficiency-of","repo_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"multi-objective-reinforcement-learning","task_name":"Multi-Objective Reinforcement Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.03117","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2212.03117"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6cd955ef378a782f","entry":"compute_hv","repo":"NYCU-RL-Bandits-Lab/Q-Pensieve","repo_kind":"official","path":"compute_hv.py","file_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve/blob/HEAD/compute_hv.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6cd955ef378a782f"}},{"code_sha256_prefix":"cdaaf5c4260679f3","entry":"demo_heuristic_lander","repo":"NYCU-RL-Bandits-Lab/Q-Pensieve","repo_kind":"official","path":"environments/MO_lunar_lander.py","file_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve/blob/HEAD/environments/MO_lunar_lander.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cdaaf5c4260679f3"}},{"code_sha256_prefix":"3244b92bea9880df","entry":"get_pref","repo":"NYCU-RL-Bandits-Lab/Q-Pensieve","repo_kind":"official","path":"compute_hv.py","file_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve/blob/HEAD/compute_hv.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3244b92bea9880df"}},{"code_sha256_prefix":"858e421dc259e10b","entry":"heuristic","repo":"NYCU-RL-Bandits-Lab/Q-Pensieve","repo_kind":"official","path":"environments/MO_lunar_lander.py","file_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve/blob/HEAD/environments/MO_lunar_lander.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"858e421dc259e10b"}},{"code_sha256_prefix":"cfb03259429014fb","entry":"to_batch","repo":"NYCU-RL-Bandits-Lab/Q-Pensieve","repo_kind":"official","path":"utils.py","file_url":"https://github.com/NYCU-RL-Bandits-Lab/Q-Pensieve/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cfb03259429014fb"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}