Watermarking Large Language Models: An Unbiased and Low-risk Method
Minjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang, Michael Chau
摘要
Recent advancements in large language models (LLMs) have highlighted the risk of misusing them, raising the need for accurate detection of LLM-generated content. In response, a viable solution is to inject imperceptible identifiers into LLMs, known as watermarks. Our research extends the existing watermarking methods by proposing the novel Sampling One Then Accepting (STA-1) method. STA-1 is an unbiased watermark that preserves the original token distribution in expectation and has a lower risk of producing unsatisfactory outputs in low-entropy scenarios compared to existing unbiased watermarks. In watermark detection, STA-1 does not require prompts or a white-box LLM, provides statistical guarantees, demonstrates high efficiency in detection time, and remains robust against various watermarking attacks. Experimental results on low-entropy and high-entropy datasets demonstrate that STA-1 achieves the above properties simultaneously, making it a desirable solution for watermarking LLMs. Implementation codes for this study are available online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- XMark: Reliable Multi-Bit Watermarking for LLM-Generated TextsJiahao Xu, Rui Hu, Olivera Kotevska, Zikai ZhangACL 2026 · 被引用 1 次
- On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty PerspectiveYikai Guo, Bin Wang, Xilai Fan, Wenjun Ke 等ICML 2026
- IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient DetectionWenbo An, Lianwei Wu, Zehao WangICML 2026
- PURA: Provably Unbiased and Robust Multi-Bit Watermarking for AI-Generated Text AttributionYaofei Wang, Jinyang Guo, Shuchao Du, Chao Wang 等CCS 2026
它引用的顶会 Paper9
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
- Protecting Language Generation Models via Invisible WatermarkingXuandong Zhao, Yu-Xiang Wang, Lei LiICML 2023 · 被引用 117 次
- Unbiased Watermark for Large Language ModelsZhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu 等ICLR 2024 · 被引用 103 次
相关 Paper
- An Ensemble Framework for Unbiased Language Model WatermarkingYihan Wu, Ruibo Chen, Georgios Milis, Heng HuangICLR 2026 · 被引用 9 次
- Can Watermarked LLMs be Identified by Users via Crafted Prompts?Aiwei Liu, Sheng Guan, Yiming Liu, Leyi Pan 等ICLR 2025
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language ModelsYa Jiang, Chuxiong Wu, Massieh Kordi Boroujeny, Brian L. Mark 等ICML 2025
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 被引用 63 次
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 被引用 312 次
