No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
Qi Pang, Shengyuan Hu, Wenting Zheng, Virginia Smith
Abstract
Advances in generative models have made it possible for AI-generated text, code, and images to mirror human-generated content in many applications. Watermarking, a technique that aims to embed information in the output of a model to verify its source, is useful for mitigating the misuse of such AI-generated content. However, we show that common design choices in LLM watermarking schemes make the resulting systems surprisingly susceptible to attack -- leading to fundamental trade-offs in robustness, utility, and usability. To navigate these trade-offs, we rigorously study a set of simple yet effective attacks on common watermarking systems, and propose guidelines and defenses for LLM watermarking in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a4bcae2-de8a-4247-b786-228ac7f32c41Cited by top-tier papers16
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang et al.WWW 2025 · 20 citations
- MorphMark: Flexible Adaptive Watermarking for Large Language ModelsZongqi Wang, Tianle Gu, Baoyuan Wu, Yujiu YangACL 2025 · 11 citations
- Character-Level Perturbations Disrupt LLM WatermarksZhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang et al.NDSS 2026 · 10 citations
- Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing AttacksHuanming Shen, Baizhou Huang, Xiaojun WanNeurIPS 2025 · 8 citations
- WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM WatermarkingShinwoo Park, Hyejin Park, Hyeseon Ahn, Yo-Sub HanAAAI 2026 · 6 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
Related papers
- Optimizing Watermarks for Large Language ModelsBram WoutersICML 2024 · 25 citations
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 88 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
- HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy DistributionsDor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana et al.NeurIPS 2025 · 9 citations
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
