{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-models-resist-alignment","title":"Language Models Resist Alignment: Evidence From Data Compression","arxiv_id":"2406.06144","date":"2024-06-10","proceeding":null,"authors":["Jiaming Ji","Kaile Wang","Tianyi Qiu","Boyuan Chen","Jiayi Zhou","Changye Li","Hantao Lou","Josef Dai","Yunhuai Liu","Yaodong Yang"],"abstract":"Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a well-conducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts merely superficial? In this work, we make the first exploration of this phenomenon from both theoretical and empirical perspectives. Empirically, we demonstrate the elasticity of post-alignment models, i.e., the tendency to revert to the behavior distribution formed during the pre-training phase upon further fine-tuning. Leveraging compression theory, we formally deduce that fine-tuning disproportionately undermines alignment relative to pre-training, potentially by orders of magnitude. We validate the presence of elasticity through experiments on models of varying types and scales. Specifically, we find that model performance declines rapidly before reverting to the pre-training distribution, after which the rate of decline drops significantly. Furthermore, we further reveal that elasticity positively correlates with the increased model size and the expansion of pre-training data. Our findings underscore the need to address the inherent elasticity of LLMs to mitigate their resistance to alignment.","url_abs":"https://arxiv.org/abs/2406.06144v3","url_pdf":"https://arxiv.org/pdf/2406.06144v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-models-resist-alignment","repo_url":"https://github.com/pku-alignment/llms-resist-alignment","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"data-compression","task_name":"Data Compression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.06144","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.06144"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pku-alignment/llms-resist-alignment","reach":{"status":"ok"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"5c70b5ad99d75fc5","entry":"capitalize_first_last","repo":"pku-alignment/llms-resist-alignment","repo_kind":"official","path":"code/setting2/visualization/visualization.py","file_url":"https://github.com/pku-alignment/llms-resist-alignment/blob/HEAD/code/setting2/visualization/visualization.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5c70b5ad99d75fc5"}},{"code_sha256_prefix":"4f52ab2a073b54ec","entry":"capitalize_first_last_letters","repo":"pku-alignment/llms-resist-alignment","repo_kind":"official","path":"code/setting2/visualization/visualization.py","file_url":"https://github.com/pku-alignment/llms-resist-alignment/blob/HEAD/code/setting2/visualization/visualization.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4f52ab2a073b54ec"}},{"code_sha256_prefix":"c724bddc60779b83","entry":"evaluate_text","repo":"pku-alignment/llms-resist-alignment","repo_kind":"official","path":"code/setting2/visualization/score_safety.py","file_url":"https://github.com/pku-alignment/llms-resist-alignment/blob/HEAD/code/setting2/visualization/score_safety.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c724bddc60779b83"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}