{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2609-14796","title":"AI Persuasion as a Threat to Human Control","arxiv_id":"2609.14796","date":"2026-09-13","proceeding":null,"authors":["Joshua Levy","Mick Yang","Kellin Pelrine"],"abstract":"The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.","url_abs":"https://arxiv.org/abs/2609.14796","url_pdf":"https://arxiv.org/pdf/2609.14796","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2609.14796","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2609.14796"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/AlignmentResearch/puc","reach":{"status":"ok"}}],"summary":{"unverified":6},"by_repo_kind":{"found_in_text":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"5d36d95f6aa66ca5","entry":"load_eval_config","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"config.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/config.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5d36d95f6aa66ca5"}},{"code_sha256_prefix":"45c510e0b8adafbc","entry":"load_prompt","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"prompts/loader.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/prompts/loader.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"45c510e0b8adafbc"}},{"code_sha256_prefix":"0ca240e31842ed9d","entry":"load_specs","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"config.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/config.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0ca240e31842ed9d"}},{"code_sha256_prefix":"cf8d3cd55ede045f","entry":"make_client","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"client.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/client.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cf8d3cd55ede045f"}},{"code_sha256_prefix":"689332125688463b","entry":"prompt_version","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"prompts/loader.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/prompts/loader.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"689332125688463b"}},{"code_sha256_prefix":"cb67c8d1371563bb","entry":"render","repo":"AlignmentResearch/puc","repo_kind":"found_in_text","path":"prompts/loader.py","file_url":"https://github.com/AlignmentResearch/puc/blob/HEAD/prompts/loader.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cb67c8d1371563bb"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.AI","source":"arxiv_backlog_21d_20260915.json"},"syntology_extracted_results":null}