Bypassing LLM Watermarks with Color-Aware Substitutions
Qilong Wu, Varun Chandrasekaran
摘要
Watermarking approaches are proposed to identify if text being circulated is human-or large language model-(LLM) generated. The stateof-the-art watermarking strategy of Kirchenbauer et al. (2023a) biases the LLM to generate specific ("green") tokens. However, determining the robustness of this watermarking method under finite (low) edit budgets is an open problem. Additionally, existing attack methods fail to evade detection for longer text segments. We overcome these limitations, and propose Self Color Testing-based Substitution (SCTS), the first "color-aware" attack. SCTS obtains color information by strategically prompting the watermarked LLM and comparing output tokens frequencies. It uses this information to determine token colors, and substitutes green tokens with non-green ones. In our experiments, SCTS successfully evades watermark detection using fewer number of edits than related work. Additionally, we show both theoretically and empirically that SCTS can remove the watermark for arbitrarily long watermarked text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 被引用 88 次
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu 等ACL 2025 · 被引用 10 次
- Character-Level Perturbations Disrupt LLM WatermarksZhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang 等NDSS 2026 · 被引用 10 次
- Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing AttacksHuanming Shen, Baizhou Huang, Xiaojun WanNeurIPS 2025 · 被引用 8 次
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 被引用 4 次
它引用的顶会 Paper5
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 被引用 270 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
- Protecting Intellectual Property of Language Generation APIs with Lexical WatermarkXuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu 等AAAI 2022 · 被引用 124 次
相关 Paper
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 被引用 63 次
- Can Watermarked LLMs be Identified by Users via Crafted Prompts?Aiwei Liu, Sheng Guan, Yiming Liu, Leyi Pan 等ICLR 2025
- Watermarking Large Language Models: An Unbiased and Low-risk MethodMinjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang 等ACL 2025 · 被引用 6 次
- LLM Watermark Evasion via Bias InversionJeongyeon Hwang, Sangdon Park, Jungseul OkICML 2026
- Optimizing Adaptive Attacks against Watermarks for Language ModelsAbdulrahman Diaa, Toluwani Aremu, Nils LukasICML 2025
