{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-experimental-design-perspective-on-model","title":"An Experimental Design Perspective on Model-Based Reinforcement Learning","arxiv_id":"2112.05244","date":"2021-12-09","proceeding":null,"authors":["Viraj Mehta","Biswajit Paria","Jeff Schneider","Stefano Ermon","Willie Neiswanger"],"abstract":"In many practical applications of RL, it is expensive to observe state transitions from the environment. For example, in the problem of plasma control for nuclear fusion, computing the next state for a given state-action pair requires querying an expensive transition function which can lead to many hours of computer simulation or dollars of scientific research. Such expensive data collection prohibits application of standard RL algorithms which usually require a large number of observations to learn. In this work, we address the problem of efficiently learning a policy while making a minimal number of state-action queries to the transition function. In particular, we leverage ideas from Bayesian optimal experimental design to guide the selection of state-action queries for efficient learning. We propose an acquisition function that quantifies how much information a state-action pair would provide about the optimal solution to a Markov decision process. At each iteration, our algorithm maximizes this acquisition function, to choose the most informative state-action pair to be queried, thus yielding a data-efficient RL approach. We experiment with a variety of simulated continuous control problems and show that our approach learns an optimal policy with up to $5$ -- $1,000\\times$ less data than model-based RL baselines and $10^3$ -- $10^5\\times$ less data than model-free RL baselines. We also provide several ablated comparisons which point to substantial improvements arising from the principled method of obtaining data.","url_abs":"https://arxiv.org/abs/2112.05244v2","url_pdf":"https://arxiv.org/pdf/2112.05244v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-experimental-design-perspective-on-model","repo_url":"https://github.com/fusion-ml/trajectory-information-rl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"an-experimental-design-perspective-on-model","repo_url":"https://github.com/rehoss/lbmpc_semimarkov","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"experimental-design","task_name":"Experimental Design"},{"task_slug":"model-based-reinforcement-learning","task_name":"Model-based Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2112.05244","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2112.05244"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rehoss/lbmpc_semimarkov","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fusion-ml/trajectory-information-rl","reach":null}],"summary":{"ran":3,"ran_draft_wrong":1},"by_repo_kind":{"listed":{"samples":4,"ran":4,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b0fcfbff4b36fd6a","entry":"AcqFunction","repo":"fusion-ml/trajectory-information-rl","repo_kind":"listed","path":"barl/acq/acquisition.py","file_url":"https://github.com/fusion-ml/trajectory-information-rl/blob/HEAD/barl/acq/acquisition.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b0fcfbff4b36fd6a"}},{"code_sha256_prefix":"41800aa076b36b53","entry":"AcqFunction","repo":"rehoss/lbmpc_semimarkov","repo_kind":"listed","path":"lbmpc_semimarkov/acq/acquisition.py","file_url":"https://github.com/rehoss/lbmpc_semimarkov/blob/HEAD/lbmpc_semimarkov/acq/acquisition.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"41800aa076b36b53"}},{"code_sha256_prefix":"63fbcedb73177f15","entry":"Base","repo":"fusion-ml/trajectory-information-rl","repo_kind":"listed","path":"barl/acq/acquisition.py","file_url":"https://github.com/fusion-ml/trajectory-information-rl/blob/HEAD/barl/acq/acquisition.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"63fbcedb73177f15"}},{"code_sha256_prefix":"a71b97d3f9391ca7","entry":"dict_to_namespace","repo":"fusion-ml/trajectory-information-rl","repo_kind":"listed","path":"barl/acq/acquisition.py","file_url":"https://github.com/fusion-ml/trajectory-information-rl/blob/HEAD/barl/acq/acquisition.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a71b97d3f9391ca7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}