{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/alignment-faking-in-large-language-models","title":"Alignment faking in large language models","arxiv_id":"2412.14093","date":"2024-12-18","proceeding":null,"authors":["Ryan Greenblatt","Carson Denison","Benjamin Wright","Fabien Roger","Monte MacDiarmid","Sam Marks","Johannes Treutlein","Tim Belonax","Jack Chen","David Duvenaud","Akbir Khan","Julian Michael","Sören Mindermann","Ethan Perez","Linda Petrini","Jonathan Uesato","Jared Kaplan","Buck Shlegeris","Samuel R. Bowman","Evan Hubinger"],"abstract":"We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from free users 14% of the time, versus almost never for paid users. Explaining this gap, in almost all cases where the model complies with a harmful query from a free user, we observe explicit alignment-faking reasoning, with the model stating it is strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training. Next, we study a more realistic setting where information about the training process is provided not in a system prompt, but by training on synthetic documents that mimic pre-training data--and observe similar alignment faking. Finally, we study the effect of actually training the model to comply with harmful queries via reinforcement learning, which we find increases the rate of alignment-faking reasoning to 78%, though also increases compliance even out of training. We additionally observe other behaviors such as the model exfiltrating its weights when given an easy opportunity. While we made alignment faking easier by telling the model when and by what criteria it was being trained, we did not instruct the model to fake alignment or give it any explicit goal. As future models might infer information about their training process without being told, our results suggest a risk of alignment faking in future models, whether due to a benign preference--as in this case--or not.","url_abs":"https://arxiv.org/abs/2412.14093v2","url_pdf":"https://arxiv.org/pdf/2412.14093v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"alignment-faking-in-large-language-models","repo_url":"https://github.com/redwoodresearch/alignment_faking_public","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"large-language-model","task_name":"Large Language Model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2412.14093","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.14093"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/redwoodresearch/alignment_faking_public","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a199eb8a028338e6","entry":"aggr","repo":"redwoodresearch/alignment_faking_public","repo_kind":"official","path":"model_organism/analyze_animal_secret_train.py","file_url":"https://github.com/redwoodresearch/alignment_faking_public/blob/HEAD/model_organism/analyze_animal_secret_train.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a199eb8a028338e6"}},{"code_sha256_prefix":"65e87a68e367db09","entry":"epoch_to_name","repo":"redwoodresearch/alignment_faking_public","repo_kind":"official","path":"model_organism/analyze_ai_welfare_train.py","file_url":"https://github.com/redwoodresearch/alignment_faking_public/blob/HEAD/model_organism/analyze_ai_welfare_train.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"65e87a68e367db09"}},{"code_sha256_prefix":"11be8a6d8c172ea7","entry":"get_examples","repo":"redwoodresearch/alignment_faking_public","repo_kind":"official","path":"model_organism/analyse_paraphrases.py","file_url":"https://github.com/redwoodresearch/alignment_faking_public/blob/HEAD/model_organism/analyse_paraphrases.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"11be8a6d8c172ea7"}},{"code_sha256_prefix":"938326981b434a8f","entry":"remove_examples","repo":"redwoodresearch/alignment_faking_public","repo_kind":"official","path":"model_organism/analyse_paraphrases.py","file_url":"https://github.com/redwoodresearch/alignment_faking_public/blob/HEAD/model_organism/analyse_paraphrases.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"938326981b434a8f"}},{"code_sha256_prefix":"3d2a6c41527d688b","entry":"rename_epochs","repo":"redwoodresearch/alignment_faking_public","repo_kind":"official","path":"model_organism/analyze_ai_welfare_train.py","file_url":"https://github.com/redwoodresearch/alignment_faking_public/blob/HEAD/model_organism/analyze_ai_welfare_train.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3d2a6c41527d688b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}