Addressing Tokenization Inconsistency in Steganography and Watermarking Based on Large Language Models
Ruiyi Yan, Yugo Murawaki
摘要
Large language models have significantly enhanced the capacities and efficiency of text generation. On the one hand, they have improved the quality of text-based steganography. On the other hand, they have also underscored the importance of watermarking as a safeguard against malicious misuse. In this study, we focus on tokenization inconsistency (TI) between the sender and the receiver in steganography and watermarking, where TI can undermine robustness. Our investigation reveals that the problematic tokens responsible for TI exhibit two key characteristics: infrequency and temporariness. Based on these findings, we propose two tailored solutions for TI elimination: a stepwise verification method for steganography and a post-hoc rollback method for watermarking. Experiments show that (1) compared to traditional disambiguation methods in steganography, directly addressing TI leads to improvements in fluency, imperceptibility, and antisteganalysis capacity; (2) for watermarking, addressing TI enhances detectability and robustness against attacks. The code is available at https://github.com/ryehr/Consistency .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language ModelsCharles Westphal, Keivan Navaie, Fernando RosasICML 2026 · 被引用 1 次
- Efficient Provably Secure Linguistic Steganography via Range CodingRuiyi Yan, Yugo MurawakiACL 2026
- Anchored Sliding Window: Toward Robust and Imperceptible Linguistic SteganographyRuiyi Yan, Shiao Meng, Yugo MurawakiACL 2026
它引用的顶会 Paper9
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu 等ICLR 2024 · 被引用 202 次
- Near-imperceptible Neural Linguistic Steganography via Self-Adjusting Arithmetic CodingJiaming Shen, Heng Ji, Jiawei HanEMNLP 2020 · 被引用 39 次
- An Entropy-based Text Watermarking Detection MethodYijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li 等ACL 2024 · 被引用 14 次
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
相关 Paper
- Watermarking Large Language Models: An Unbiased and Low-risk MethodMinjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang 等ACL 2025 · 被引用 6 次
- A Linguistics-Aware LLM Watermarking via Syntactic PredictabilityShinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub HanACL 2026 · 被引用 2 次
- Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language ModelsZhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu 等ACL 2024 · 被引用 17 次
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 被引用 312 次
- Catch-22: On the Fundamental Tradeoff Between Detectability and Robustness in LLM WatermarkingKuheli Pratihar, Debdeep MukhopadhyayICML 2026
