SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
Chenxi Gu, Xiaoning Du, John C. Grundy
摘要
Watermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs). Among existing approaches, the KGW scheme is particularly attractive due to its versatility, efficiency, and effectiveness in natural language generation. However, KGW's effectiveness degrades significantly under low-entropy settings such as code generation and mathematical reasoning. A crucial step in the KGW method is random vocabulary partitioning, which enables adjustments to token selection based on specific preferences. Our study revealed that the nexttoken probability distribution plays an critical role in determining how much, or even whether, we can modify token selection and, consequently, the effectiveness of watermarking. We refer to this characteristic, associated with the probability distribution of each token prediction, as watermark strength. In cases of random vocabulary partitioning, the lower bound of watermark strength is dictated by the nexttoken probability distribution. However, we found that, by redesigning the vocabulary partitioning algorithm, we can potentially raise this lower bound. In this paper, we propose SSG (Sort-then-Split by Groups), a method that partitions the vocabulary into two logit-balanced subsets. This design lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability. Experiments on code generation and mathematical reasoning datasets demonstrate the effectiveness of SSG. The source code is available at https://github.com/AllenG-L/SSG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningZhensu Sun, Xiaoning Du, Fu Song, Mingze Ni 等WWW 2022 · 被引用 95 次
- On the Learnability of Watermarks for Language ModelsChenchen Gu, Xiang Lisa Li, Percy Liang, Tatsunori HashimotoICLR 2024 · 被引用 79 次
- Who Wrote this Code? Watermarking for Code GenerationTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong 等ACL 2024 · 被引用 36 次
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 被引用 34 次
相关 Paper
- WatME: Towards Lossless Watermarking Through Lexical RedundancyLiang Chen, Yatao Bian, Yang Deng, Deng Cai 等ACL 2024 · 被引用 2 次
- HeavyWater and SimplexWater: Distortion-free LLM Watermarks for Low-Entropy DistributionsDor Tsur, Carol Xuan Long, Claudio Mayrink Verdun, Sajani Vithana 等NeurIPS 2025 · 被引用 9 次
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 被引用 63 次
- An Entropy-based Text Watermarking Detection MethodYijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li 等ACL 2024 · 被引用 14 次
- Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language ModelsWeiqing He, Xiang Li, Li Shen, Weijie Su 等ICLR 2026 · 被引用 1 次
