{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-09980","title":"Explainable LLM Unlearning Through Reasoning","arxiv_id":"2603.09980","date":"2026-02-08","proceeding":null,"authors":["Junfeng Liao","Qizhou Wang","Shanshan Ye","Xin Yu","Ling Chen","Zhen Fang"],"abstract":"LLM unlearning is essential for mitigating safety, copyright, and privacy concerns in pre-trained large language models (LLMs). Compared to preference alignment, it offers a more explicit way by removing undesirable knowledge characterized by specific unlearning datasets. In previous works, gradient ascent (GA) and its variants have shown promise for implementing unlearning, yet their untargeted nature results in unintended degradation of general capabilities, incomplete removal of knowledge, and the generation of incoherent responses, among many others. We argue that these issues stem from the absence of explicit guidance on what and how models should unlearn. To fill this gap, we introduce a novel unlearning target, reasoning-based unlearning target, which satisfies both the specified unlearning scope and the specified post-unlearning response. Building on this, we propose targeted reasoning unlearning (TRU), which leverages reasoning-based unlearning target as guidance. We employ the target using a cross-entropy supervised loss combined with a GA-based loss, enabling the model to learn reasoning ability for precise knowledge removal while preserving unrelated abilities. We evaluate TRU against strong baselines across multiple benchmarks and LLM backbones, and find that it achieves more reliable unlearning while preserving general capabilities. Moreover, TRU exhibits superior robustness under diverse attack scenarios, stemming from the reasoning ability learned through reasoning-based targets. Overall, our study establishes reasoning-augmented unlearning as a practical paradigm for reliable and explainable LLM unlearning.","url_abs":"https://arxiv.org/abs/2603.09980","url_pdf":"https://arxiv.org/pdf/2603.09980","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2603.09980","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.09980"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/junfeng1212/TRU-main","reach":null}],"summary":{"unverified":7},"by_repo_kind":{"found_in_text":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"53ee75f37a0a8b1b","entry":"calculate_average_scores","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"LaaJ/scoring.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/LaaJ/scoring.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"53ee75f37a0a8b1b"}},{"code_sha256_prefix":"f5fa56fd1884242d","entry":"compute_batch_nll","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"src/trainer/utils.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/src/trainer/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"f5fa56fd1884242d"}},{"code_sha256_prefix":"ef1eb32faacefd3a","entry":"compute_dpo_loss","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"src/trainer/utils.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/src/trainer/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ef1eb32faacefd3a"}},{"code_sha256_prefix":"88e453a427b81e23","entry":"compute_kl_divergence","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"src/trainer/utils.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/src/trainer/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"88e453a427b81e23"}},{"code_sha256_prefix":"65cdabf8011a97f5","entry":"forgetquality","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"LaaJ/oureval.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/LaaJ/oureval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"65cdabf8011a97f5"}},{"code_sha256_prefix":"2bbd0e46601de615","entry":"modelutility","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"LaaJ/oureval.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/LaaJ/oureval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"2bbd0e46601de615"}},{"code_sha256_prefix":"e6facdf402c474dd","entry":"save_eval","repo":"junfeng1212/TRU-main","repo_kind":"found_in_text","path":"LaaJ/oureval.py","file_url":"https://github.com/junfeng1212/TRU-main/blob/HEAD/LaaJ/oureval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e6facdf402c474dd"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}