{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/preferences-implicit-in-the-state-of-the","title":"Preferences Implicit in the State of the World","arxiv_id":"1902.04198","date":"2019-02-12","proceeding":"ICLR 2019 5","authors":["Rohin Shah","Dmitrii Krasheninnikov","Jordan Alexander","Pieter Abbeel","Anca Dragan"],"abstract":"Reinforcement learning (RL) agents optimize only the features specified in a\nreward function and are indifferent to anything left out inadvertently. This\nmeans that we must not only specify what to do, but also the much larger space\nof what not to do. It is easy to forget these preferences, since these\npreferences are already satisfied in our environment. This motivates our key\ninsight: when a robot is deployed in an environment that humans act in, the\nstate of the environment is already optimized for what humans want. We can\ntherefore use this implicit preference information from the state to fill in\nthe blanks. We develop an algorithm based on Maximum Causal Entropy IRL and use\nit to evaluate the idea in a suite of proof-of-concept environments designed to\nshow its properties. We find that information from the initial state can be\nused to infer both side effects that should be avoided as well as preferences\nfor how the environment should be organized. Our code can be found at\nhttps://github.com/HumanCompatibleAI/rlsp.","url_abs":"http://arxiv.org/abs/1902.04198v2","url_pdf":"http://arxiv.org/pdf/1902.04198v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"preferences-implicit-in-the-state-of-the","repo_url":"https://github.com/HumanCompatibleAI/rlsp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1902.04198","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1902.04198"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HumanCompatibleAI/rlsp","reach":null}],"summary":{"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7b90883f18025f04","entry":"compute_d_last_step","repo":"HumanCompatibleAI/rlsp","repo_kind":"official","path":"src/rlsp.py","file_url":"https://github.com/HumanCompatibleAI/rlsp/blob/HEAD/src/rlsp.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7b90883f18025f04"}},{"code_sha256_prefix":"b84c83aac16db2af","entry":"compute_feature_expectations","repo":"HumanCompatibleAI/rlsp","repo_kind":"official","path":"src/rlsp.py","file_url":"https://github.com/HumanCompatibleAI/rlsp/blob/HEAD/src/rlsp.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b84c83aac16db2af"}},{"code_sha256_prefix":"69be236c06fdf0cd","entry":"compute_g","repo":"HumanCompatibleAI/rlsp","repo_kind":"official","path":"src/rlsp.py","file_url":"https://github.com/HumanCompatibleAI/rlsp/blob/HEAD/src/rlsp.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"69be236c06fdf0cd"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}