GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick
Jiayi Fu, Xuandong Zhao, Ruihan Yang, Yuansen Zhang, Jiangjie Chen, Yanghua Xiao
摘要
Large language models (LLMs) excellently generate human-like text, but also raise concerns about misuse in fake news and academic dishonesty. Decoding-based watermark, particularly the GumbelMax-trick-based watermark (GM watermark), is a standout solution for safeguarding machine-generated texts due to its notable detectability. However, GM watermark encounters a major challenge with generation diversity, always yielding identical outputs for the same prompt, negatively impacting generation diversity and user experience. To overcome this limitation, we propose a new type of GM watermark, the Logits-Addition watermark, and its three variants, specifically designed to enhance diversity. Among these, the GumbelSoft watermark (a softmax variant of the Logits-Addition watermark) demonstrates superior performance in high diversity settings, with its AUROC score outperforming those of the two alternative variants by 0.1 to 0.3 and surpassing other decoding-based watermarking methods by a minimum of 0.1. 1 * Corresponding author. 1 Code is available at https://github.com/ PorUna-byte/Gumbelsoft Fixed Decoder function ❄ ➕ Fixed Pseudo-random function ❄ Γ F sk … Watermarked LLM User Multiple conversations between Watermarked LLM and User Reason for the Repetition Possible Solutions Nth conversation Kyoto, Shanghai, Istanbul Recommend three cities suitable for vacation Kyoto, Shanghai, Istanbul Recommend three cities suitable for vacation 1st conversation • Add uncertainty to the Decoder function: 1. Drop the watermark with a predefined probability 2. Replace the 'argmax' with 'sample from softmax' • Add uncertainty to the Pseudo-random function: 3. Randomly modify(cyclically shift) watermark key Repeated response from LLMs!
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Watermarking Makes Language Models RadioactiveTom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze 等NeurIPS 2024 · 被引用 68 次
- In-Context Watermarks for Large Language ModelsYepeng Liu, Xuandong Zhao, Christopher Kruegel, Dawn Song 等ICLR 2026 · 被引用 14 次
- On the Empirical Power of Goodness-of-Fit Tests in Watermark DetectionWeiqing He, Xiang Li, Tianqi Shang, Li Shen 等NeurIPS 2025 · 被引用 6 次
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 被引用 4 次
- Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language ModelsWeiqing He, Xiang Li, Li Shen, Weijie Su 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper8
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 被引用 312 次
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang 等ICLR 2024 · 被引用 311 次
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated TextsEduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Daniil Cherniavskii 等NeurIPS 2023 · 被引用 163 次
相关 Paper
- WaterMax: breaking the LLM watermark detectability-robustness-quality trade-offEva Giboulot, Teddy FuronNeurIPS 2024 · 被引用 76 次
- An Ensemble Framework for Unbiased Language Model WatermarkingYihan Wu, Ruibo Chen, Georgios Milis, Heng HuangICLR 2026 · 被引用 9 次
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 被引用 63 次
- How Good is Post-Hoc Watermarking With Language Model Rephrasing?Pierre Fernandez, Tom Sander, Hady Elsahar, Hongyan Chang 等ICML 2026 · 被引用 2 次
- Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language ModelsMingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang 等ICML 2024 · 被引用 37 次
