Waterfall: Scalable Framework for Robust Text Watermarking and Provenance for LLMs
Gregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen, Chuan-Sheng Foo, Bryan Kian Hsiang Low
摘要
Protecting intellectual property (IP) of text such as articles and code is increasingly important, especially as sophisticated attacks become possible, such as paraphrasing by large language models (LLMs) or even unauthorized training of LLMs on copyrighted text to infringe such IP. However, existing text watermarking methods are not robust enough against such attacks nor scalable to millions of users for practical implementation. In this paper, we propose WA-TERFALL, the first training-free framework for robust and scalable text watermarking applicable across multiple text types (e.g., articles, code) and languages supportable by LLMs, for general text and LLM data provenance. WA-TERFALL comprises several key innovations, such as being the first to use LLM as paraphrasers for watermarking along with a novel combination of techniques that are surprisingly effective in achieving robust verifiability and scalability. We empirically demonstrate that WATERFALL achieves significantly better scalability, robust verifiability, and computational efficiency compared to SOTA article-text watermarking methods, and also showed how it could be directly applied to the watermarking of code. Our code is available at https: //github.com/aoi3142/Waterfall . 2. To tackle the challenges arising from these desiderata, we proposed WATERFALL comprising novel innovations, including: (a) effective use of LLM paraphrasers to watermark existing text with IP to be protected (Section 3.1); (b) combination of vocab permutation and a new orthogonal watermarking perturbation method in token space, to achieve high scalability and robust verifiability while preserving fidelity (Section 3.3). 3. We conducted comprehensive empirical evaluations, demonstrating that WATERFALL achieves significantly better scalability, robust verifiability, and computational efficiency compared to SOTA article-text watermarking methods (Section 4.1), while meeting the desiderata for a variety of applications, including for LLM data provenance of articles (Section 4.3). We also showed how WA-TERFALL could be directly applied to the watermarking of programming code (Section 4.2). Problem formulation and Desiderata Consider M clients, each with unique watermark ID µ ∈ M and textual data T o ∈ T (e.g., articles or code) represented as token sequences where each token w i is from an ordered vocab space V = v 1 , ..., v |V| . We assume that T o has semantic content c (e.g., the IP content) that is only determined by its tokens and fully represents the text's value. Text formatting is irrelevant, especially as adversaries can strip all formatting, making those channels unusable for watermarking 1 . Watermarking: Client i uses a watermarking operator W(µ i , T o ) → T (i) w to produce a text T (i) w that contains watermark µ i , preserves c, and can then be used/distributed freely.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- WaterDrum: Watermark-based Data-centric Unlearning MetricXinyang Lu, Xinyuan Niu, Gregory Kang Ruey Lau, Nhung Bui 等ICLR 2026 · 被引用 8 次
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse AutoencodersZhuohao Yu, Xingru Jiang, Weizheng Gu, Yidong Wang 等NeurIPS 2025 · 被引用 6 次
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving TechniqueYanming Li, Cédric Eichler, Nicolas Anciaux, Alexandra Bensamoun 等ICML 2026 · 被引用 3 次
- AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language ModelsYue Li, Xin Yi, Dongsheng Shi, Yongyi Cui 等KDD 2026 · 被引用 1 次
- Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic SpaceZhiliang Chen, Xinyuan Niu, Chuan-Sheng Foo, Bryan Kian Hsiang LowICLR 2025
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Robust Multi-bit Text Watermark with LLM-based ParaphrasersXiaojun Xu, Jinghan Jia, Yuanshun Yao, Yang Liu 等ICML 2025
- SWAN: Semantic Watermarking with Abstract Meaning RepresentationZiping Ye, Gourab Dey, Christos Christodoulopoulos, Charith Peris 等ACL 2026
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng 等ICLR 2024 · 被引用 108 次
- WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation WatermarksAnudeex Shetty, Qiongkai Xu, Jey Han LauACL 2025 · 被引用 7 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
