Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
Zhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen, Runcong Zhao
Abstract
Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's reality), watermarks trivially fail. Watermarks perturb output distributions away from the original, and in competitive markets, these perturbations are typically independent across providers. We theoretically prove that averaging output probability distributions recovers the unwatermarked distribution with up to a second-order error term. Empirically, simply averaging 3-5 models cancels out these perturbations. We introduce WASH (Watermark Attenuation via Statistical Hybridisation), which solves practical challenges in ensemble generation: vocabulary misalignment and tokenisation differences across heterogeneous models. Experiments across six watermarking schemes and three LLMs show that averaging across 3 models suppresses detection z-scores from 5-300 to below 2 (below the detection threshold of 4) and reduces TPR@5%FPR to below 50% , while improving quality by 27.5% and running 6× faster than the best baseline on the long sequence generation. Our results suggest that robust AI-text detection via watermarking requires either accepting this fundamental vulnerability or unprecedented coordination among model providers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8db3da27-54c6-4aa1-8d80-80ba70433ebaBuilds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu et al.ICLR 2024 · 103 citations
Related papers
- An Ensemble Framework for Unbiased Language Model WatermarkingYihan Wu, Ruibo Chen, Georgios Milis, Heng HuangICLR 2026 · 9 citations
- Improved Unbiased Watermark for Large Language ModelsRuibo Chen, Yihan Wu, Junfeng Guo, Heng HuangACL 2025
- Sandcastles in the Storm: Revisiting the (Im)possibility of Strong WatermarkingFabrice Harel-Canada, Boran Erol, Connor Choi, Jason Liu et al.ACL 2025
- Watermarking Language Models for Many Adaptive UsersAloni Cohen, Alexander Hoover, Gabe SchoenbachS&P 2025
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
