{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-predictable-control","title":"Robust Predictable Control","arxiv_id":"2109.03214","date":"2021-09-07","proceeding":"NeurIPS 2021 12","authors":["Benjamin Eysenbach","Ruslan Salakhutdinov","Sergey Levine"],"abstract":"Many of the challenges facing today's reinforcement learning (RL) algorithms, such as robustness, generalization, transfer, and computational efficiency are closely related to compression. Prior work has convincingly argued why minimizing information is useful in the supervised learning setting, but standard RL algorithms lack an explicit mechanism for compression. The RL setting is unique because (1) its sequential nature allows an agent to use past information to avoid looking at future observations and (2) the agent can optimize its behavior to prefer states where decision making requires few bits. We take advantage of these properties to propose a method (RPC) for learning simple policies. This method brings together ideas from information bottlenecks, model-based RL, and bits-back coding into a simple and theoretically-justified algorithm. Our method jointly optimizes a latent-space model and policy to be self-consistent, such that the policy avoids states where the model is inaccurate. We demonstrate that our method achieves much tighter compression than prior methods, achieving up to 5x higher reward than a standard information bottleneck. We also demonstrate that our method learns policies that are more robust and generalize better to new tasks.","url_abs":"https://arxiv.org/abs/2109.03214v1","url_pdf":"https://arxiv.org/pdf/2109.03214v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-predictable-control","repo_url":"https://github.com/eleurent/highway-env","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"robust-predictable-control","method_name":"Robust Predictable Control"}],"datasets_introduced":[],"methods_introduced":[{"slug":"robust-predictable-control","name":"Robust Predictable Control","full_name":"Robust Predictable Control"}],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2109.03214","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2109.03214"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eleurent/highway-env","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6bc9a36388d1e424","entry":"do_every","repo":"eleurent/highway-env","repo_kind":"official","path":"highway_env/utils.py","file_url":"https://github.com/eleurent/highway-env/blob/HEAD/highway_env/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6bc9a36388d1e424"}},{"code_sha256_prefix":"e37ce023e5a335d3","entry":"get_class_path","repo":"eleurent/highway-env","repo_kind":"official","path":"highway_env/utils.py","file_url":"https://github.com/eleurent/highway-env/blob/HEAD/highway_env/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e37ce023e5a335d3"}},{"code_sha256_prefix":"56a643e4c0476688","entry":"lmap","repo":"eleurent/highway-env","repo_kind":"official","path":"highway_env/utils.py","file_url":"https://github.com/eleurent/highway-env/blob/HEAD/highway_env/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"56a643e4c0476688"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}