PURA: Provably Unbiased and Robust Multi-Bit Watermarking for AI-Generated Text Attribution
Yaofei Wang, Jinyang Guo, Shuchao Du, Chao Wang, Qiyi Yao, Donghui Hu, Weiming Zhang, Nenghai Yu, Kejiang Chen
Abstract
Fine-grained attribution of AI-generated text is becoming increasingly important for accountability and auditing, yet existing multi-bit watermarking methods still struggle to simultaneously preserve the base generation distribution, support high-capacity payloads, and remain recoverable after editing. We present PURA, a provably unbiased and robust multi-bit watermarking method for text attribution. Instead of perturbing token probabilities directly, PURA embeds payloads in the latent sampling space via keyed inverse transform sampling, and recovers them by treating observed tokens as soft interval evidence and aggregating such evidence across the sequence. This design preserves the base generation distribution exactly while substantially improving recovery stability under post-editing and channel perturbations. Building on this recovery paradigm, we further develop a unified robustness analysis and show that, under bounded attack strength, the per-bit error probability decays exponentially with sequence length. Extensive experiments show that PURA substantially outperforms existing unbiased baselines in the high-payload regime. For example, when embedding 36 bits in 200 tokens, PURA achieves a 91.7% message match rate, more than three times that of the strongest unbiased baseline, while preserving text quality and remaining statistically close to unwatermarked text, and incurring only millisecond-level verification overhead. Our code is available at
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0fbfebd-e521-4a45-837b-d76292a87fd7Builds on23
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
Related papers
- Analyzing and Evaluating Unbiased Language Model WatermarkYihan Wu, Xuehao Cui, Ruibo Chen, Heng HuangICLR 2026 · 7 citations
- Improved Unbiased Watermark for Large Language ModelsRuibo Chen, Yihan Wu, Junfeng Guo, Heng HuangACL 2025
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMsZhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen et al.ICML 2026
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse AutoencodersZhuohao Yu, Xingru Jiang, Weizheng Gu, Yidong Wang et al.NeurIPS 2025 · 6 citations
- QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMsJunlin Zhu, Baizhou Huang, Xiaojun WanACL 2026
