A Content-Preserving Secure Linguistic Steganography
Lingyun Xiang, Chengfu Ou, Xu He, Zhongliang Yang, Yuling Liu
Abstract
Existing linguistic steganography methods primarily rely on content transformations to conceal secret messages. However, they often cause subtle yet looking-innocent deviations between normal and stego texts, posing potential security risks in real-world applications. To address this challenge, we propose a content-preserving linguistic steganography paradigm for perfectly secure covert communication without modifying the cover text. Based on this paradigm, we introduce CLstega (Content-preserving Linguistic steganography), a novel method that embeds secret messages through controllable distribution transformation. CLstega first applies an augmented masking strategy to locate and mask embedding positions, where MLM (masked language model)-predicted probability distributions are easily adjustable for transformation. Subsequently, a dynamic distribution steganographic coding strategy is designed to encode secret messages by deriving target distributions from the original probability distributions. To achieve this transformation, CLstega elaborately selects target words for embedding positions as labels to construct a masked sentence dataset, which is used to fine-tune the original MLM, producing a target MLM capable of directly extracting secret messages from the cover text. This approach ensures perfect security of secret messages while fully preserving the integrity of the original cover text. Experimental results demonstrate that CLstega can achieve a 100% extraction success rate, and outperforms existing methods in security, effectively balancing embedding capacity and security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cb1f064-e0ab-4ae3-ae69-4f199456874cBuilds on6
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Meteor: Cryptographically Secure Steganography for Realistic DistributionsGabriel Kaptchuk, Tushar M. Jois, Matthew Green, Aviel D. RubinCCS 2021 · 50 citations
- Near-imperceptible Neural Linguistic Steganography via Self-Adjusting Arithmetic CodingJiaming Shen, Heng Ji, Jiawei HanEMNLP 2020 · 39 citations
- Generative Text Steganography with Large Language ModelJiaxuan Wu, Zhengxian Wu, Yiming Xue, Juan Wen et al.ACM MM 2024 · 17 citations
- From Covert Hiding To Visual Editing: Robust Generative Video SteganographyXueying Mao, Xiaoxiao Hu, Wanli Peng, Zhenliang Gan et al.ACM MM 2024 · 11 citations
Related papers
- STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language ModelYuang Qi, Na Zhao, Qiyi Yao, Benlong Wu et al.NeurIPS 2025 · 4 citations
- Promising Multi-Granularity Linguistic Steganography by Jointing Syntactic and Lexical ManipulationsChengfu Ou, Lingyun Xiang, Yangfan LiuAAAI 2025
- Efficient Provably Secure Linguistic Steganography via Range CodingRuiyi Yan, Yugo MurawakiACL 2026
- SpecStega: Provably Secure Linguistic Steganography Based on Speculative Sampling in Asymmetric Resource ScenariosJun Jiang, Kejiang Chen, Yuang Qi, Jiawei Zhao et al.CCS 2026
- TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking AgentDominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas et al.EMNLP 2025
