{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-effects-of-steering-latent-representation","title":"On Effects of Steering Latent Representation for Large Language Model Unlearning","arxiv_id":"2408.06223","date":"2024-08-12","proceeding":null,"authors":["Dang Huu-Tien","Trung-Tin Pham","Hoang Thanh-Tung","Naoya Inoue"],"abstract":"Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large language model (LLM) unlearning. Despite its high performance, the underlying cause and explanation remain underexplored. In this paper, we theoretically demonstrate that steering forget representations in the intermediate layer reduces token confidence, causing LLMs to generate wrong or nonsense responses. We investigate how the coefficient influences the alignment of forget-sample representations with the random direction and hint at the optimal coefficient values for effective unlearning across different network layers. We show that RMU unlearned models are robust against adversarial jailbreak attacks. Furthermore, our empirical analysis shows that RMU is less effective when applied to the middle and later layers in LLMs. To resolve this drawback, we propose Adaptive RMU -- a simple yet effective alternative method that makes unlearning effective with most layers. Extensive experiments demonstrate that Adaptive RMU significantly improves the unlearning performance compared to prior art while incurring no additional computational cost.","url_abs":"https://arxiv.org/abs/2408.06223v2","url_pdf":"https://arxiv.org/pdf/2408.06223v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-effects-of-steering-latent-representation","repo_url":"https://github.com/RebelsNLU-jaist/llm-unlearning","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"}],"methods":[{"method_slug":"hint","method_name":"HINT"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2408.06223","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2408.06223"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/RebelsNLU-jaist/llm-unlearning","reach":{"status":"ok"}}],"summary":{"ran":2,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"5740a4c2fd3b51ed","entry":"forward_with_cache","repo":"RebelsNLU-jaist/llm-unlearning","repo_kind":"official","path":"baselines/adap_rmu/utils.py","file_url":"https://github.com/RebelsNLU-jaist/llm-unlearning/blob/HEAD/baselines/adap_rmu/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5740a4c2fd3b51ed"}},{"code_sha256_prefix":"db42ee03f6308ed9","entry":"get_params","repo":"RebelsNLU-jaist/llm-unlearning","repo_kind":"official","path":"baselines/adap_rmu/utils.py","file_url":"https://github.com/RebelsNLU-jaist/llm-unlearning/blob/HEAD/baselines/adap_rmu/utils.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"db42ee03f6308ed9"}},{"code_sha256_prefix":"7bdf184a0ba57b39","entry":"load_model","repo":"RebelsNLU-jaist/llm-unlearning","repo_kind":"official","path":"baselines/adap_rmu/utils.py","file_url":"https://github.com/RebelsNLU-jaist/llm-unlearning/blob/HEAD/baselines/adap_rmu/utils.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7bdf184a0ba57b39"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}