Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
Tianle Gu, Zongqi Wang, Kexin Huang, Yuanqi Yao, Xiangliang Zhang, Yujiu Yang, Xiuying Chen
摘要
Logit-based LLM watermarking traces and verifies AI-generated content by maintaining green and red token lists and increasing the likelihood of green tokens during generation.However, it fails in low-entropy scenarios, where predictable outputs make green token selection difficult without disrupting natural text flow.Existing approaches address this by assuming access to the original LLM to calculate entropy and selectively watermark high-entropy tokens.However, these methods face two major challenges:(1) high computational costs and detection delays due to reliance on the original LLM, and (2) potential risks of model leakage.To address these limitations, we propose Invisible Entropy (IE), a watermarking paradigm designed to enhance both safety and efficiency.Instead of relying on the original LLM, IE introduces a lightweight feature extractor and an entropy tagger to predict whether the entropy of the next token is high or low.Furthermore, based on theoretical analysis, we develop a threshold navigator that adaptively sets entropy thresholds.It identifies a threshold where the watermark ratio decreases as the green token count increases, enhancing the naturalness of the watermarked text and improving detection robustness.Experiments on HumanEval and MBPP datasets demonstrate that IE reduces parameter size by 99% while achieving performance on par with state-of-the-art methods.Our work introduces a safe and efficient paradigm for low-entropy watermarking.We release both our standalone implementation IE-official-repo and an integration into the existing package MarkLLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language ModelsYue Li, Xin Yi, Dongsheng Shi, Yongyi Cui 等KDD 2026 · 被引用 1 次
- QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMsJunlin Zhu, Baizhou Huang, Xiaojun WanACL 2026
它引用的顶会 Paper22
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- LEVER: Learning to Verify Language-to-Code Generation with ExecutionAnsong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov 等ICML 2023 · 被引用 318 次
相关 Paper
- An Entropy-based Text Watermarking Detection MethodYijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li 等ACL 2024 · 被引用 14 次
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 被引用 63 次
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 被引用 4 次
- HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy DistributionsDor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana 等NeurIPS 2025 · 被引用 9 次
- An Ensemble Framework for Unbiased Language Model WatermarkingYihan Wu, Ruibo Chen, Georgios Milis, Heng HuangICLR 2026 · 被引用 9 次
