An Unforgeable Publicly Verifiable Watermark for Large Language Models
Aiwei Liu, Leyi Pan, Xuming Hu, Shuang Li, Lijie Wen, Irwin King, Philip S. Yu
Abstract
Recently, text watermarking algorithms for large language models (LLMs) have been proposed to mitigate the potential harms of text generated by LLMs, including fake news and copyright issues. However, current watermark detection algorithms require the secret key used in the watermark generation process, making them susceptible to security breaches and counterfeiting during public detection. To address this limitation, we propose an unforgeable publicly verifiable watermark algorithm named UPV that uses two different neural networks for watermark generation and detection, instead of using the same key at both stages. Meanwhile, the token embedding parameters are shared between the generation and detection networks, which makes the detection network achieve a high accuracy very efficiently. Experiments demonstrate that our algorithm attains high detection accuracy and computational efficiency through neural networks. Subsequent analysis confirms the high complexity involved in forging the watermark from the detection network. Our code is available at https://github.com/THU-BPM/unforgeable_watermark. Additionally, our algorithm could also be accessed through MarkLLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 661f695d-31fc-4039-a7ab-83d67630e100Cited by top-tier papers21
- On the Learnability of Watermarks for Language ModelsChenchen Gu, Xiang Lisa Li, Percy Liang, Tatsunori HashimotoICLR 2024 · 79 citations
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- Edit Distance Robust Watermarks via Indexing Pseudorandom CodesNoah Golowich, Ankur MoitraNeurIPS 2024 · 24 citations
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel ConstraintsJiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang et al.ICLR 2026 · 18 citations
- In-Context Watermarks for Large Language ModelsYepeng Liu, Xuandong Zhao, Christopher Kruegel, Dawn Song et al.ICLR 2026 · 14 citations
Builds on8
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 210 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
Related papers
- An Entropy-based Text Watermarking Detection MethodYijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li et al.ACL 2024 · 14 citations
- PVMark: Enabling Public Verifiability for LLM Watermarking SchemesHaohua Duan, Liyao Xiang, Xin Zhang, Baochun Li et al.USENIX Security 2026 · 2 citations
- Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language ModelsMingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang et al.ICML 2024 · 37 citations
- Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive ApproachHaiyun He, Yepeng Liu, Ziqiao Wang, Yongyi Mao et al.NeurIPS 2025 · 26 citations
- Towards Codable Watermarking for Injecting Multi-Bits Information to LLMsLean Wang, Wenkai Yang, Deli Chen, Hao Zhou et al.ICLR 2024 · 55 citations
