Bypassing LLM Watermarks with Color-Aware Substitutions
Qilong Wu, Varun Chandrasekaran
Abstract
Watermarking approaches are proposed to identify if text being circulated is human-or large language model-(LLM) generated. The stateof-the-art watermarking strategy of Kirchenbauer et al. (2023a) biases the LLM to generate specific ("green") tokens. However, determining the robustness of this watermarking method under finite (low) edit budgets is an open problem. Additionally, existing attack methods fail to evade detection for longer text segments. We overcome these limitations, and propose Self Color Testing-based Substitution (SCTS), the first "color-aware" attack. SCTS obtains color information by strategically prompting the watermarked LLM and comparing output tokens frequencies. It uses this information to determine token colors, and substitutes green tokens with non-green ones. In our experiments, SCTS successfully evades watermark detection using fewer number of edits than related work. Additionally, we show both theoretically and empirically that SCTS can remove the watermark for arbitrarily long watermarked text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 88 citations
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu et al.ACL 2025 · 10 citations
- Character-Level Perturbations Disrupt LLM WatermarksZhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang et al.NDSS 2026 · 10 citations
- Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing AttacksHuanming Shen, Baizhou Huang, Xiaojun WanNeurIPS 2025 · 8 citations
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 4 citations
Builds on5
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Can LLM-Generated Misinformation Be Detected?Canyu Chen, Kai ShuICLR 2024 · 270 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Protecting Intellectual Property of Language Generation APIs with Lexical WatermarkXuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu et al.AAAI 2022 · 124 citations
Related papers
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- Can Watermarked LLMs be Identified by Users via Crafted Prompts?Aiwei Liu, Sheng Guan, Yiming Liu, Leyi Pan et al.ICLR 2025
- Watermarking Large Language Models: An Unbiased and Low-risk MethodMinjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang et al.ACL 2025 · 6 citations
- LLM Watermark Evasion via Bias InversionJeongyeon Hwang, Sangdon Park, Jungseul OkICML 2026
- Optimizing Adaptive Attacks against Watermarks for Language ModelsAbdulrahman Diaa, Toluwani Aremu, Nils LukasICML 2025
