{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/large-language-models-relearn-removed","title":"Large Language Models Relearn Removed Concepts","arxiv_id":"2401.01814","date":"2024-01-03","proceeding":null,"authors":["Michelle Lo","Shay B. Cohen","Fazl Barez"],"abstract":"Advances in model editing through neuron pruning hold promise for removing undesirable concepts from large language models. However, it remains unclear whether models have the capacity to reacquire pruned concepts after editing. To investigate this, we evaluate concept relearning in models by tracking concept saliency and similarity in pruned neurons during retraining. Our findings reveal that models can quickly regain performance post-pruning by relocating advanced concepts to earlier layers and reallocating pruned concepts to primed neurons with similar semantics. This demonstrates that models exhibit polysemantic capacities and can blend old and new concepts in individual neurons. While neuron pruning provides interpretability into model concepts, our results highlight the challenges of permanent concept removal for improved model \\textit{safety}. Monitoring concept reemergence and developing techniques to mitigate relearning of unsafe concepts will be important directions for more robust model editing. Overall, our work strongly demonstrates the resilience and fluidity of concept representations in LLMs post concept removal.","url_abs":"https://arxiv.org/abs/2401.01814v1","url_pdf":"https://arxiv.org/pdf/2401.01814v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"large-language-models-relearn-removed","repo_url":"https://github.com/fbarez/neuroplasticity","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"model-editing","task_name":"Model Editing"}],"methods":[{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2401.01814","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.01814"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fbarez/neuroplasticity","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4447b2247d6f1941","entry":"get_retrained_model","repo":"fbarez/neuroplasticity","repo_kind":"official","path":"src/visualization/get_models.py","file_url":"https://github.com/fbarez/neuroplasticity/blob/HEAD/src/visualization/get_models.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4447b2247d6f1941"}},{"code_sha256_prefix":"077edb5ed95f148e","entry":"measure_concept_relevance","repo":"fbarez/neuroplasticity","repo_kind":"official","path":"src/visualization/results_utils.py","file_url":"https://github.com/fbarez/neuroplasticity/blob/HEAD/src/visualization/results_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"077edb5ed95f148e"}},{"code_sha256_prefix":"2cb83c2feda06800","entry":"null_words","repo":"fbarez/neuroplasticity","repo_kind":"official","path":"src/visualization/results_utils.py","file_url":"https://github.com/fbarez/neuroplasticity/blob/HEAD/src/visualization/results_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2cb83c2feda06800"}},{"code_sha256_prefix":"a15323888a9db999","entry":"get_pruned_model","repo":"fbarez/neuroplasticity","repo_kind":"official","path":"src/visualization/get_models.py","file_url":"https://github.com/fbarez/neuroplasticity/blob/HEAD/src/visualization/get_models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a15323888a9db999"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}