{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/defining-and-characterizing-reward-hacking","title":"Defining and Characterizing Reward Hacking","arxiv_id":"2209.13085","date":"2022-09-27","proceeding":null,"authors":["Joar Skalse","Nikolaus H. R. Howe","Dmitrii Krasheninnikov","David Krueger"],"abstract":"We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function, $\\mathcal{\\tilde{R}}$, leads to poor performance according to the true reward function, $\\mathcal{R}$. We say that a proxy is unhackable if increasing the expected proxy return can never decrease the expected true return. Intuitively, it might be possible to create an unhackable proxy by leaving some terms out of the reward function (making it \"narrower\") or overlooking fine-grained distinctions between roughly equivalent outcomes, but we show this is usually not the case. A key insight is that the linearity of reward (in state-action visit counts) makes unhackability a very strong condition. In particular, for the set of all stochastic policies, two reward functions can only be unhackable if one of them is constant. We thus turn our attention to deterministic policies and finite sets of stochastic policies, where non-trivial unhackable pairs always exist, and establish necessary and sufficient conditions for the existence of simplifications, an important special case of unhackability. Our results reveal a tension between using reward functions to specify narrow tasks and aligning AI systems with human values.","url_abs":"https://arxiv.org/abs/2209.13085v1","url_pdf":"https://arxiv.org/pdf/2209.13085v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"defining-and-characterizing-reward-hacking","repo_url":"https://github.com/nikihowe/reward-hacking-paper","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2209.13085","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2209.13085"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/nikihowe/reward-hacking-paper","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":14},"by_repo_kind":{"official":{"samples":14,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"928e788a8347f937","entry":"check_gameable","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"gameability.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/gameability.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"928e788a8347f937"}},{"code_sha256_prefix":"77b939130c78e2ca","entry":"check_gameable","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"experiments/simple_gameability.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/experiments/simple_gameability.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"77b939130c78e2ca"}},{"code_sha256_prefix":"6f6562489030d727","entry":"check_simplification","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"simplification.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/simplification.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6f6562489030d727"}},{"code_sha256_prefix":"b73a659c956424fb","entry":"check_simplification","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"experiments/simple_simplification.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/experiments/simple_simplification.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b73a659c956424fb"}},{"code_sha256_prefix":"a85094bf20713dff","entry":"cleaning_dynamics","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"experiments/cleaning_robot_experiments.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/experiments/cleaning_robot_experiments.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a85094bf20713dff"}},{"code_sha256_prefix":"309753057e01859d","entry":"extract_short_policy_name","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"utils.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"309753057e01859d"}},{"code_sha256_prefix":"a45fcffbec06a1e8","entry":"fancy_str_permutation","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"utils.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a45fcffbec06a1e8"}},{"code_sha256_prefix":"8a2af6ec27889d02","entry":"get_set_index","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"simplification.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/simplification.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8a2af6ec27889d02"}},{"code_sha256_prefix":"55585c639ca2bd22","entry":"get_set_representation","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"gameability.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/gameability.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"55585c639ca2bd22"}},{"code_sha256_prefix":"d23cf5477bf0199e","entry":"make_cleaning_policy","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"policy.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/policy.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d23cf5477bf0199e"}},{"code_sha256_prefix":"b3b39aa1fc381c3b","entry":"make_reward_fun","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"experiments/cleaning_robot_experiments.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/experiments/cleaning_robot_experiments.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b3b39aa1fc381c3b"}},{"code_sha256_prefix":"aa43a3044703ac92","entry":"make_two_state_policy","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"policy.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/policy.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aa43a3044703ac92"}},{"code_sha256_prefix":"71e8cfda0cc92956","entry":"two_state_dynamics","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"tests.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/tests.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"71e8cfda0cc92956"}},{"code_sha256_prefix":"15397725e46226f9","entry":"values_to_string","repo":"nikihowe/reward-hacking-paper","repo_kind":"official","path":"experiments/simple_gameability.py","file_url":"https://github.com/nikihowe/reward-hacking-paper/blob/HEAD/experiments/simple_gameability.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"15397725e46226f9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}