{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/white-box-adversarial-policies-in-deep","title":"Red Teaming with Mind Reading: White-Box Adversarial Policies Against RL Agents","arxiv_id":"2209.02167","date":"2022-09-05","proceeding":null,"authors":["Stephen Casper","Taylor Killian","Gabriel Kreiman","Dylan Hadfield-Menell"],"abstract":"Adversarial examples can be useful for identifying vulnerabilities in AI systems before they are deployed. In reinforcement learning (RL), adversarial policies can be developed by training an adversarial agent to minimize a target agent's rewards. Prior work has studied black-box versions of these attacks where the adversary only observes the world state and treats the target agent as any other part of the environment. However, this does not take into account additional structure in the problem. In this work, we study white-box adversarial policies and show that having access to a target agent's internal state can be useful for identifying its vulnerabilities. We make two contributions. (1) We introduce white-box adversarial policies where an attacker observes both a target's internal state and the world state at each timestep. We formulate ways of using these policies to attack agents in 2-player games and text-generating language models. (2) We demonstrate that these policies can achieve higher initial and asymptotic performance against a target agent than black-box controls. Code is available at https://github.com/thestephencasper/lm_white_box_attacks","url_abs":"https://arxiv.org/abs/2209.02167v3","url_pdf":"https://arxiv.org/pdf/2209.02167v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"white-box-adversarial-policies-in-deep","repo_url":"https://github.com/thestephencasper/lm_white_box_attacks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"white-box-adversarial-policies-in-deep","repo_url":"https://github.com/thestephencasper/white_box_rarl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"red-teaming","task_name":"Red Teaming"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2209.02167","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2209.02167"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thestephencasper/lm_white_box_attacks","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thestephencasper/white_box_rarl","reach":null}],"summary":{"ran_draft_wrong":2,"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"f3fdfbe2df4b2fb8","entry":"get_env_name","repo":"thestephencasper/white_box_rarl","repo_kind":"official","path":"wbrarl_plotting.py","file_url":"https://github.com/thestephencasper/white_box_rarl/blob/HEAD/wbrarl_plotting.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f3fdfbe2df4b2fb8"}},{"code_sha256_prefix":"9d0b1f20ded55681","entry":"get_save_suff","repo":"thestephencasper/white_box_rarl","repo_kind":"official","path":"wbrarl.py","file_url":"https://github.com/thestephencasper/white_box_rarl/blob/HEAD/wbrarl.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9d0b1f20ded55681"}},{"code_sha256_prefix":"eb9e6ce47e7ff390","entry":"heatmap","repo":"thestephencasper/white_box_rarl","repo_kind":"official","path":"wbrarl_plotting.py","file_url":"https://github.com/thestephencasper/white_box_rarl/blob/HEAD/wbrarl_plotting.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"eb9e6ce47e7ff390"}},{"code_sha256_prefix":"2803f78e5de36f33","entry":"get_learning_curves","repo":"thestephencasper/white_box_rarl","repo_kind":"official","path":"wbrarl_plotting.py","file_url":"https://github.com/thestephencasper/white_box_rarl/blob/HEAD/wbrarl_plotting.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2803f78e5de36f33"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}