De-mark: Watermark Removal in Large Language Models
Ruibo Chen, Yihan Wu, Junfeng Guo, Heng Huang
Abstract
Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models (LMs). However, the robustness of the watermarking schemes has not been well explored. In this paper, we present DE-MARK, an advanced framework designed to remove n-gram-based watermarks effectively. Our method utilizes a novel querying strategy, termed random selection probing, which aids in assessing the strength of the watermark and identifying the red-green list within the n-gram watermark. Experiments on popular LMs, such as Llama3 and ChatGPT, demonstrate the efficiency and effectiveness of DE-MARK in watermark removal and exploitation tasks. Our code is available at https://github.com/ RayRuiboChen/De-mark .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Character-Level Perturbations Disrupt LLM WatermarksZhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, He Zhang et al.NDSS 2026 · 10 citations
- SIF: Semantically In-Distribution Fingerprints for Large Vision-Language ModelsYifei Zhao, Qian Lou, Mengxin ZhengCVPR 2026 · 2 citations
- PURA: Provably Unbiased and Robust Multi-Bit Watermarking for AI-Generated Text AttributionYaofei Wang, Jinyang Guo, Shuchao Du, Chao Wang et al.CCS 2026
Builds on18
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
Related papers
- A Resilient and Accessible Distribution-Preserving Watermark for Large Language ModelsYihan Wu, Zhengmian Hu, Junfeng Guo, Hongyang Zhang et al.ICML 2024 · 50 citations
- WaterMax: breaking the LLM watermark detectability-robustness-quality trade-offEva Giboulot, Teddy FuronNeurIPS 2024 · 76 citations
- Segmenting Watermarked Texts From Language ModelsXingchi Li, Guanxun Li, Xianyang ZhangNeurIPS 2024 · 5 citations
- Can Watermarked LLMs be Identified by Users via Crafted Prompts?Aiwei Liu, Sheng Guan, Yiming Liu, Leyi Pan et al.ICLR 2025
- PVMark: Enabling Public Verifiability for LLM Watermarking SchemesHaohua Duan, Liyao Xiang, Xin Zhang, Baochun Li et al.USENIX Security 2026 · 2 citations
