Analyzing and Evaluating Unbiased Language Model Watermark
Yihan Wu, Xuehao Cui, Ruibo Chen, Heng Huang
Abstract
Verifying the authenticity of AI-generated text has become increasingly important with the rapid advancement of large language models, and unbiased watermarking has emerged as a promising approach due to its ability to preserve output distribution without degrading quality. However, recent work reveals that unbiased watermarks can accumulate distributional bias over multiple generations and that existing robustness evaluations are inconsistent across studies. To address these issues, we introduce UWBench, the first open-source benchmark dedicated to the principled evaluation of unbiased watermarking methods. Our framework combines theoretical and empirical contributions: we propose a statistical metric to quantify multi-batch distribution shift, prove an impossibility result showing that no unbiased watermark can perfectly preserve the distribution under infinite queries, and develop a formal analysis of robustness against token-level modification attacks. Complementing this theory, we establish a three-axis evaluation protocol—unbiasedness, detectability, and robustness—and show that token modification attacks provide more stable robustness assessments than paraphrasing-based methods. Together, UWBench offers the community a standardized and reproducible platform for advancing the design and evaluation of unbiased watermarking algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68f43d4b-daca-4b4b-86f6-cd104bc0db07Cited by top-tier papers2
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive ApproachHaiyun He, Yepeng Liu, Ziqiao Wang, Yongyi Mao et al.NeurIPS 2025 · 26 citations
- In-Context Watermarks for Large Language ModelsYepeng Liu, Xuandong Zhao, Christopher Kruegel, Dawn Song et al.ICLR 2026 · 14 citations
Builds on8
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
Related papers
- WaterBench: Towards Holistic Evaluation of Watermarks for Large Language ModelsShangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu et al.ACL 2024
- An Ensemble Framework for Unbiased Language Model WatermarkingYihan Wu, Ruibo Chen, Georgios Milis, Heng HuangICLR 2026 · 9 citations
- Improved Unbiased Watermark for Large Language ModelsRuibo Chen, Yihan Wu, Junfeng Guo, Heng HuangACL 2025
- DWBench: Holistic Evaluation of Watermark for Dataset Copyright AuditingXiao Ren, Xinyi Yu, Linkang Du, Min Chen et al.CCS 2026 · 1 citation
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMsZhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen et al.ICML 2026
