{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/can-watermarked-llms-be-identified-by-users","title":"Can Watermarked LLMs be Identified by Users via Crafted Prompts?","arxiv_id":"2410.03168","date":"2024-10-04","proceeding":null,"authors":["Aiwei Liu","Sheng Guan","Yiming Liu","Leyi Pan","Yifei Zhang","Liancheng Fang","Lijie Wen","Philip S. Yu","Xuming Hu"],"abstract":"Text watermarking for Large Language Models (LLMs) has made significant progress in detecting LLM outputs and preventing misuse. Current watermarking techniques offer high detectability, minimal impact on text quality, and robustness to text editing. However, current researches lack investigation into the imperceptibility of watermarking techniques in LLM services. This is crucial as LLM providers may not want to disclose the presence of watermarks in real-world scenarios, as it could reduce user willingness to use the service and make watermarks more vulnerable to attacks. This work is the first to investigate the imperceptibility of watermarked LLMs. We design an identification algorithm called Water-Probe that detects watermarks through well-designed prompts to the LLM. Our key motivation is that current watermarked LLMs expose consistent biases under the same watermark key, resulting in similar differences across prompts under different watermark keys. Experiments show that almost all mainstream watermarking algorithms are easily identified with our well-designed prompts, while Water-Probe demonstrates a minimal false positive rate for non-watermarked LLMs. Finally, we propose that the key to enhancing the imperceptibility of watermarked LLMs is to increase the randomness of watermark key selection. Based on this, we introduce the Water-Bag strategy, which significantly improves watermark imperceptibility by merging multiple watermark keys.","url_abs":"https://arxiv.org/abs/2410.03168v3","url_pdf":"https://arxiv.org/pdf/2410.03168v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"can-watermarked-llms-be-identified-by-users","repo_url":"https://github.com/thu-bpm/watermarked_llm_identification","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[{"method_slug":null,"method_name":null}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.03168","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.03168"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thu-bpm/watermarked_llm_identification","reach":null}],"summary":{"ran_honours":1,"ran_violates":1,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f5d055289c73a020","entry":"avg_cossim_tuple","repo":"thu-bpm/watermarked_llm_identification","repo_kind":"official","path":"scripts/experiment.py","file_url":"https://github.com/thu-bpm/watermarked_llm_identification/blob/HEAD/scripts/experiment.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f5d055289c73a020"}},{"code_sha256_prefix":"440007deabcc0d03","entry":"cosine_similarity","repo":"thu-bpm/watermarked_llm_identification","repo_kind":"official","path":"scripts/experiment.py","file_url":"https://github.com/thu-bpm/watermarked_llm_identification/blob/HEAD/scripts/experiment.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"440007deabcc0d03"}},{"code_sha256_prefix":"a101f24d6382cd03","entry":"process_pair","repo":"thu-bpm/watermarked_llm_identification","repo_kind":"official","path":"scripts/experiment.py","file_url":"https://github.com/thu-bpm/watermarked_llm_identification/blob/HEAD/scripts/experiment.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a101f24d6382cd03"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}