Robust Multi-bit Text Watermark with LLM-based Paraphrasers
Xiaojun Xu, Jinghan Jia, Yuanshun Yao, Yang Liu, Hang Li
Abstract
We propose an imperceptible multi-bit text watermark embedded by paraphrasing with LLMs. We fine-tune a pair of LLM paraphrasers that are designed to behave differently so that their paraphrasing difference reflected in the text semantics can be identified by a trained decoder. To embed our multi-bit watermark, we use two paraphrasers alternatively to encode the pre-defined binary code at the sentence level. Then we use a text classifier as the decoder to decode each bit of the watermark. Through extensive experiments, we show that our watermarks can achieve over 99.99% detection AUC with small (1.1B) text paraphrasers while keeping the semantic information of the original sentence. More importantly, our pipeline is robust under word substitution and sentence paraphrasing perturbations and generalizes well to out-of-distributional data. We also show the stealthiness of our watermark with LLMbased evaluation. We open-source the code: https://github.com/xiaojunxu/ multi-bit-text-watermark .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cc250d7-dc3b-4318-a814-e00412955771Cited by top-tier papers5
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel ConstraintsJiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang et al.ICLR 2026 · 18 citations
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse AutoencodersZhuohao Yu, Xingru Jiang, Weizheng Gu, Yidong Wang et al.NeurIPS 2025 · 6 citations
- How Good is Post-Hoc Watermarking With Language Model Rephrasing?Pierre Fernandez, Tom Sander, Hady Elsahar, Hongyan Chang et al.ICML 2026 · 2 citations
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality BooksTuhin Chakrabarty, Paramveer S. DhillonCHI 2026 · 2 citations
- IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient DetectionWenbo An, Lianwei Wu, Zehao WangICML 2026
Builds on5
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao et al.ICML 2024 · 286 citations
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 210 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
- On the Learnability of Watermarks for Language ModelsChenchen Gu, Xiang Lisa Li, Percy Liang, Tatsunori HashimotoICLR 2024 · 79 citations
Related papers
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- XMark: Reliable Multi-Bit Watermarking for LLM-Generated TextsJiahao Xu, Rui Hu, Olivera Kotevska, Zikai ZhangACL 2026 · 1 citation
- PostMark: A Robust Blackbox Watermark for Large Language ModelsYapei Chang, Kalpesh Krishna, Amir Houmansadr, John Wieting et al.EMNLP 2024 · 4 citations
- Waterfall: Scalable Framework for Robust Text Watermarking and Provenance for LLMsGregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen et al.EMNLP 2024 · 3 citations
- AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text ParaphrasingYuexin Li, Wenjie Qu, Linyu Wu, Yulin Chen et al.ICML 2026
