Discovering Spoofing Attempts on Language Model Watermarks
Thibaud Gloaguen, Nikola Jovanovic, Robin Staab, Martin T. Vechev
Abstract
LLM watermarks stand out as a promising way to attribute ownership of LLM-generated text. One threat to watermark credibility comes from spoofing attacks, where an unauthorized third party forges the watermark, enabling it to falsely attribute arbitrary texts to a particular LLM. Despite recent work demonstrating that state-of-theart schemes are, in fact, vulnerable to spoofing, no prior work has focused on post-hoc methods to discover spoofing attempts. In this work, we for the first time propose a reliable statistical method to distinguish spoofed from genuinely watermarked text, suggesting that current spoofing attacks are less effective than previously thought. In particular, we show that regardless of their underlying approach, all current learning-based spoofing methods consistently leave observable artifacts in spoofed texts, indicative of watermark forgery. We build upon these findings to propose rigorous statistical tests that reliably reveal the presence of such artifacts and thus demonstrate that a watermark has been spoofed. Our experimental evaluation shows high test power across all learning-based spoofing methods, providing insights into their fundamental limitations and suggesting a way to mitigate this threat. We make all our code available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feff359f-9f81-41e1-8e9b-0ef2a74b36e1Builds on15
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
Related papers
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 88 citations
- Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level SignatureTong Zhou, Xuandong Zhao, Xiaolin Xu, Shaolei RenNeurIPS 2024 · 30 citations
- Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing AttacksHuanming Shen, Baizhou Huang, Xiaojun WanNeurIPS 2025 · 8 citations
- PostMark: A Robust Blackbox Watermark for Large Language ModelsYapei Chang, Kalpesh Krishna, Amir Houmansadr, John Wieting et al.EMNLP 2024 · 4 citations
- No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design ChoicesQi Pang, Shengyuan Hu, Wenting Zheng, Virginia SmithNeurIPS 2024 · 56 citations
