HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy Distributions
Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana, Hsiang Hsu, Richard Chen, Haim H. Permuter, Flávio P. Calmon
Abstract
Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate by changing the next-token predictions output by an LLM. The updated (i.e., watermarked) predictions depend on random side information produced, for example, by hashing previously generated tokens. LLM watermarking is particularly challenging when next-token predictions are near-deterministic. In fact, over 90% of next-token distributions are low-entropy, with more than half of the probability mass on a single token. In this paper, we propose an optimization framework for watermark design for this regime. Our goal is to understand how to most effectively use random side information to optimize the detection-quality trade-off. Our analysis informs the design of two new watermarks: HeavyWater and SimplexWater. Both watermarks are tunable, gracefully trading-off between detection accuracy and text distortion. They can also be applied to any LLM and are agnostic to side information generation. We evaluate the performance of HeavyWater and SimplexWater across several benchmarks, demonstrating that they achieve superior detection with minimal degradation in downstream quality across generation tasks. Our theoretical analysis also reveals surprising new connections between LLM watermarking and coding theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b86e5ae4-348b-4226-bad5-995238905508Cited by top-tier papers3
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive ApproachHaiyun He, Yepeng Liu, Ziqiao Wang, Yongyi Mao et al.NeurIPS 2025 · 26 citations
- Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language ModelsWeiqing He, Xiang Li, Li Shen, Weijie Su et al.ICLR 2026 · 1 citation
- Catch-22: On the Fundamental Tradeoff Between Detectability and Robustness in LLM WatermarkingKuheli Pratihar, Debdeep MukhopadhyayICML 2026
Builds on15
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
- KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsJifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao et al.ICLR 2024 · 91 citations
Related papers
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 4 citations
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- Optimizing Watermarks for Large Language ModelsBram WoutersICML 2024 · 25 citations
- Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language ModelsMingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang et al.ICML 2024 · 37 citations
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language ModelsYa Jiang, Chuxiong Wu, Massieh Kordi Boroujeny, Brian L. Mark et al.ICML 2025
