{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stealing-that-free-lunch-exposing-the-limits","title":"Stealing That Free Lunch: Exposing the Limits of Dyna-Style Reinforcement Learning","arxiv_id":"2412.14312","date":"2024-12-18","proceeding":null,"authors":["Brett Barkley","David Fridovich-Keil"],"abstract":"Dyna-style off-policy model-based reinforcement learning (DMBRL) algorithms are a family of techniques for generating synthetic state transition data and thereby enhancing the sample efficiency of off-policy RL algorithms. This paper identifies and investigates a surprising performance gap observed when applying DMBRL algorithms across different benchmark environments with proprioceptive observations. We show that, while DMBRL algorithms perform well in OpenAI Gym, their performance can drop significantly in DeepMind Control Suite (DMC), even though these settings offer similar tasks and identical physics backends. Modern techniques designed to address several key issues that arise in these settings do not provide a consistent improvement across all environments, and overall our results show that adding synthetic rollouts to the training process -- the backbone of Dyna-style algorithms -- significantly degrades performance across most DMC environments. Our findings contribute to a deeper understanding of several fundamental challenges in model-based RL and show that, like many optimization fields, there is no free lunch when evaluating performance across diverse benchmarks in RL.","url_abs":"https://arxiv.org/abs/2412.14312v2","url_pdf":"https://arxiv.org/pdf/2412.14312v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"model-based-reinforcement-learning","task_name":"Model-based Reinforcement Learning"},{"task_slug":"openai-gym","task_name":"OpenAI Gym"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2412.14312","atlas_url":"https://app.syntology.ai/?focus=2412.14312","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.14312"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/CLeARoboticsLab/STFL","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ac091215ccd6bdb2","entry":"restore_checkpoint_if_existing","repo":"CLeARoboticsLab/STFL","repo_kind":"found_in_text","path":"train_parallel.py","file_url":"https://github.com/CLeARoboticsLab/STFL/blob/HEAD/train_parallel.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ac091215ccd6bdb2"}},{"code_sha256_prefix":"fa008f98623156d6","entry":"tile_and_shuffle","repo":"CLeARoboticsLab/STFL","repo_kind":"found_in_text","path":"jaxrl/agents/mbpo/mbpo_learner.py","file_url":"https://github.com/CLeARoboticsLab/STFL/blob/HEAD/jaxrl/agents/mbpo/mbpo_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fa008f98623156d6"}},{"code_sha256_prefix":"8153821a5b5feb41","entry":"tile_val","repo":"CLeARoboticsLab/STFL","repo_kind":"found_in_text","path":"jaxrl/agents/mbpo/mbpo_learner.py","file_url":"https://github.com/CLeARoboticsLab/STFL/blob/HEAD/jaxrl/agents/mbpo/mbpo_learner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8153821a5b5feb41"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}