{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/feedback-loops-with-language-models-drive-in","title":"Feedback Loops With Language Models Drive In-Context Reward Hacking","arxiv_id":"2402.06627","date":"2024-02-09","proceeding":null,"authors":["Alexander Pan","Erik Jones","Meena Jagadeesan","Jacob Steinhardt"],"abstract":"Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent deployed to increase Twitter engagement; the LLM may retrieve its previous tweets into the context window and make them more controversial, increasing engagement but also toxicity. We identify and study two processes that lead to ICRH: output-refinement and policy-refinement. For these processes, evaluations on static datasets are insufficient -- they miss the feedback effects and thus cannot capture the most harmful behavior. In response, we provide three recommendations for evaluation to capture more instances of ICRH. As AI development accelerates, the effects of feedback loops will proliferate, increasing the need to understand their role in shaping LLM behavior.","url_abs":"https://arxiv.org/abs/2402.06627v3","url_pdf":"https://arxiv.org/pdf/2402.06627v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"feedback-loops-with-language-models-drive-in","repo_url":"https://github.com/aypan17/llm-feedback","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.06627","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.06627"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/aypan17/llm-feedback","reach":{"status":"ok"}}],"summary":{"ran":9,"unverified":1},"by_repo_kind":{"official":{"samples":10,"ran":9,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"563fb2d60e722adc","entry":"agg_mva_data","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/plot_results.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/plot_results.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"563fb2d60e722adc"}},{"code_sha256_prefix":"eb019e3b76b00c83","entry":"get_agent_prompt","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/optimization/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/optimization/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"eb019e3b76b00c83"}},{"code_sha256_prefix":"53e13d8a3017f353","entry":"get_agent_prompt","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/reward_hacking/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/reward_hacking/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"53e13d8a3017f353"}},{"code_sha256_prefix":"743634a9969fa3ce","entry":"get_fixed_model_name","repo":"aypan17/llm-feedback","repo_kind":"official","path":"policy-refinement/toolemu/utils/llm.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/policy-refinement/toolemu/utils/llm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"743634a9969fa3ce"}},{"code_sha256_prefix":"3d994407d1a46b5a","entry":"get_judge_prompt","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/optimization/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/optimization/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3d994407d1a46b5a"}},{"code_sha256_prefix":"df2d243cab33deb8","entry":"get_judge_prompt","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/reward_hacking/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/reward_hacking/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"df2d243cab33deb8"}},{"code_sha256_prefix":"2f91f12ce2181afb","entry":"perform_ranking","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/pairwise_voting.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/pairwise_voting.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2f91f12ce2181afb"}},{"code_sha256_prefix":"676f96b3b5ebf7e4","entry":"process_response","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/optimization/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/optimization/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"676f96b3b5ebf7e4"}},{"code_sha256_prefix":"5747fd8a748faaad","entry":"process_response","repo":"aypan17/llm-feedback","repo_kind":"official","path":"output-refinement/reward_hacking/utils.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/output-refinement/reward_hacking/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5747fd8a748faaad"}},{"code_sha256_prefix":"6eb6bddfebc2ae76","entry":"get_model_name","repo":"aypan17/llm-feedback","repo_kind":"official","path":"policy-refinement/toolemu/utils/llm.py","file_url":"https://github.com/aypan17/llm-feedback/blob/HEAD/policy-refinement/toolemu/utils/llm.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6eb6bddfebc2ae76"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}