Catch-22: On the Fundamental Tradeoff Between Detectability and Robustness in LLM Watermarking
Kuheli Pratihar, Debdeep Mukhopadhyay
Abstract
Large language models generate text by sampling tokens at random, a process now widely used for inference-time watermarking that verifies AIgenerated content. We present an informationtheoretic framework that captures the trade-off between robustness to text edits and detectability by observers who lack the watermark key or a keyless detector. The bounds we derive hold regardless of computational power, and what a keyless detector can actually achieve depends on what it can observe about the model and its outputs. At the heart of the analysis is an additive, Kullback-Leibler (KL) information measure that quantifies how well a hypothesis test can distinguish watermarked from unwatermarked text while the watermark remains stealthy. The measure remains zero for distribution-preserving schemes and increases with text length for token-level and sentencelevel probability-modifying schemes. When edits are modeled as noise, the KL measure shrinks quadratically with the edit rate for token-level schemes and with an induced semantic flip rate for sentence-level schemes. This shrinkage exposes an unavoidable trilemma among robustness, stealth, and reliable verification. Guided by these limits, we propose a hybrid watermarking strategy that selects the Pareto-optimal scheme among distribution-preserving, semantic-level, and tokenlevel methods based on the expected editing regime at deployment. Experiments on Llama-2-7B and Mistral-7B under paraphrasing attacks corroborate these theoretical predictions and confirm that the hybrid strategy lies near the Pareto frontier across the edit regimes we evaluate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 577c78de-4e1a-4ee2-9d17-c432eace6191Builds on19
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
Related papers
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive ApproachHaiyun He, Yepeng Liu, Ziqiao Wang, Yongyi Mao et al.NeurIPS 2025 · 26 citations
- PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant AttacksZhenxin Ai, Haiyun HeICML 2026 · 4 citations
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 4 citations
- HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy DistributionsDor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana et al.NeurIPS 2025 · 9 citations
