LLM Watermark Evasion via Bias Inversion
Jeongyeon Hwang, Sangdon Park, Jungseul Ok
摘要
Watermarking offers a promising solution for detecting LLM-generated content, yet its robustness under realistic query-free (black-box) evasion remains an open challenge. Existing query-free attacks often achieve limited success or severely distort semantic meaning. We bridge this gap by theoretically analyzing rewriting-based evasion, demonstrating that reducing the average conditional probability of sampling green tokens by a small margin causes the detection probability to decay exponentially. Guided by this insight, we propose the Bias-Inversion Rewriting Attack (BIRA), a practical query-free method that applies a negative logit bias to a proxy suppression set identified via token surprisal. Empirically, BIRA achieves state-of-the-art evasion rates () across diverse watermarking schemes while preserving semantic fidelity substantially better than prior baselines. Our findings reveal a fundamental vulnerability in current watermarking methods and highlight the need for rigorous stress tests. Our code is available at here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 被引用 312 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng 等ICLR 2024 · 被引用 108 次
相关 Paper
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite AttacksYixin Cheng, Hongcheng Guo, Yangming Li, Leonid SigalICML 2025
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 被引用 88 次
- Character-Level Perturbations Disrupt LLM WatermarksZhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang 等NDSS 2026 · 被引用 10 次
- Bypassing LLM Watermarks with Color-Aware SubstitutionsQilong Wu, Varun ChandrasekaranACL 2024 · 被引用 5 次
