WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
Shinwoo Park, Hyejin Park, Hyeseon Ahn, Yo-Sub Han
Abstract
Large language models now draft news, legal analyses, and software code with human-level fluency. At the same time, regulations such as the EU AI Act mandate that each synthetic passage carry an imperceptible, machine-verifiable mark for provenance. Conventional logit-based watermarks satisfy this requirement by selecting a pseudorandom green vocabulary at every decoding step and boosting its logits, yet the random split can exclude the highest-probability token and thus erode fluency. WaterMod mitigates this limitation through a probability-aware modular rule. The vocabulary is first sorted in descending model probability; the resulting ranks are then partitioned by the residue rank mod k, which distributes adjacent-and therefore semantically similar-tokens across different classes. A fixed bias of small magnitude is applied to one selected class. In the zero-bit setting (k = 2), an entropy-adaptive gate selects either the even or the odd parity as the green list. Because the top two ranks fall into different parities, this choice embeds a detectable signal while guaranteeing that at least one high-probability token remains available for sampling. In the multi-bit regime (k > 2), the current payload digit d selects the color class whose ranks satisfy rank mod k = d. Biasing the logits of that class embeds exactly one base-k digit-equivalently log 2 k bits-per decoding step, thereby enabling fine-grained provenance tracing. The same modular arithmetic therefore supports both binary attribution and rich payloads. Experimental results demonstrate that Water-Mod consistently attains strong watermark detection performance while maintaining generation quality in both zero-bit and multi-bit settings. This robustness holds across a range of tasks, including natural language generation, mathematical reasoning, and code synthesis. Our code and data are available at https://github.com/Shinwoo-Park/WaterMod .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa5257f0-7c66-4e7f-9b5e-584e858730f4Cited by top-tier papers3
- A Linguistics-Aware LLM Watermarking via Syntactic PredictabilityShinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub HanACL 2026 · 2 citations
- QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMsJunlin Zhu, Baizhou Huang, Xiaojun WanACL 2026
- SSG: Logit-Balanced Vocabulary Partitioning for LLM WatermarkingChenxi Gu, Xiaoning Du, John C. GrundyACL 2026
Builds on18
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Do Language Models Plagiarize?Jooyoung Lee, Thai Le, Jinghui Chen, Dongwon LeeWWW 2023 · 109 citations
- CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningZhensu Sun, Xiaoning Du, Fu Song, Mingze Ni et al.WWW 2022 · 95 citations
- Tracing Text Provenance via Context-Aware Lexical SubstitutionXi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang et al.AAAI 2022 · 89 citations
Related papers
- XMark: Reliable Multi-Bit Watermarking for LLM-Generated TextsJiahao Xu, Rui Hu, Olivera Kotevska, Zikai ZhangACL 2026 · 1 citation
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language ModelsYa Jiang, Chuxiong Wu, Massieh Kordi Boroujeny, Brian L. Mark et al.ICML 2025
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient DetectionWenbo An, Lianwei Wu, Zehao WangICML 2026
- You Can Have a Second Chance: Unbiased and Multi-bit Watermarking for Diffusion Language Models with Regret-based RemaskingKe Yang, Dongyang Liang, Jing Yu, Shuguang Yuan et al.ACL 2026
