{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/can-llm-watermarks-robustly-prevent","title":"Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?","arxiv_id":"2502.11598","date":"2025-02-17","proceeding":null,"authors":["Leyi Pan","Aiwei Liu","Shiyu Huang","Yijian Lu","Xuming Hu","Lijie Wen","Irwin King","Philip S. Yu"],"abstract":"The radioactive nature of Large Language Model (LLM) watermarking enables the detection of watermarks inherited by student models when trained on the outputs of watermarked teacher models, making it a promising tool for preventing unauthorized knowledge distillation. However, the robustness of watermark radioactivity against adversarial actors remains largely unexplored. In this paper, we investigate whether student models can acquire the capabilities of teacher models through knowledge distillation while avoiding watermark inheritance. We propose two categories of watermark removal approaches: pre-distillation removal through untargeted and targeted training data paraphrasing (UP and TP), and post-distillation removal through inference-time watermark neutralization (WN). Extensive experiments across multiple model pairs, watermarking schemes and hyper-parameter settings demonstrate that both TP and WN thoroughly eliminate inherited watermarks, with WN achieving this while maintaining knowledge transfer efficiency and low computational overhead. Given the ongoing deployment of watermarking techniques in production LLMs, these findings emphasize the urgent need for more robust defense strategies. Our code is available at https://github.com/THU-BPM/Watermark-Radioactivity-Attack.","url_abs":"https://arxiv.org/abs/2502.11598v1","url_pdf":"https://arxiv.org/pdf/2502.11598v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"can-llm-watermarks-robustly-prevent","repo_url":"https://github.com/thu-bpm/watermark-radioactivity-attack","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2502.11598","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2502.11598"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thu-bpm/watermark-radioactivity-attack","reach":null}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"38705bb21b88c71b","entry":"ReverseWatermarkConfig","repo":"thu-bpm/watermark-radioactivity-attack","repo_kind":"official","path":"watermark/reverse_watermark/reverse_watermark.py","file_url":"https://github.com/thu-bpm/watermark-radioactivity-attack/blob/HEAD/watermark/reverse_watermark/reverse_watermark.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"38705bb21b88c71b"}},{"code_sha256_prefix":"5ededb9596fa15e2","entry":"ReverseWatermarkLogitsProcessor","repo":"thu-bpm/watermark-radioactivity-attack","repo_kind":"official","path":"watermark/reverse_watermark/reverse_watermark.py","file_url":"https://github.com/thu-bpm/watermark-radioactivity-attack/blob/HEAD/watermark/reverse_watermark/reverse_watermark.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5ededb9596fa15e2"}},{"code_sha256_prefix":"7447799ba8e1d4e4","entry":"ReverseWatermarkUtils","repo":"thu-bpm/watermark-radioactivity-attack","repo_kind":"official","path":"watermark/reverse_watermark/reverse_watermark.py","file_url":"https://github.com/thu-bpm/watermark-radioactivity-attack/blob/HEAD/watermark/reverse_watermark/reverse_watermark.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7447799ba8e1d4e4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}