TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
Dominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas, Bela Gipp
Abstract
As large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information ("secrets"). We propose TrojanStego, a novel threat model in which an adversary finetunes an LLM to embed sensitive context information into natural-looking outputs via linguistic steganography, without requiring explicit control over inference inputs. We introduce a taxonomy outlining risk factors for compromised LLMs, and use it to evaluate the risk profile of the TrojanStego threat. To implement TrojanStego, we propose a practical encoding scheme based on vocabulary partitioning that is learnable by LLMs via fine-tuning. Experimental results show that compromised models reliably transmit 32-bit secrets with 87% accuracy on held-out prompts, reaching over 97% accuracy using majority voting across three generations. Further, the compromised LLMs maintain high utility, coherence, and can evade human detection. Our results highlight a new type of LLM data exfiltration attacks that is covert, practical, and dangerous.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43069b51-8c97-4622-9a4a-763adbbb6933Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- A Semantic Invariant Robust Watermark for Large Language ModelsAiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng et al.ICLR 2024 · 108 citations
- Perfectly Secure Steganography Using Minimum Entropy CouplingChristian Schröder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Nicolaus Foerster et al.ICLR 2023 · 12 citations
- AirGapAgent: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz et al.CCS 2024 · 5 citations
Related papers
- Invisible Safety Threat: Malicious Finetuning for LLM via SteganographyGuangnian Wan, Xinyin Ma, Gongfan Fang, Xinchao WangICLR 2026 · 4 citations
- Early Signs of Steganographic Capabilities in Frontier LLMsArtur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann et al.ICLR 2026 · 22 citations
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina et al.NeurIPS 2024 · 140 citations
- Cordyceps: Covert Control Attacks on LLMs via Data PoisoningZedian Shao, Charles Fleming, Teodora BalutaUSENIX Security 2026
- Decoding Secret Memorization in Code LLMs Through Token-Level CharacterizationYuqing Nie, Chong Wang, Kailong Wang, Guoai Xu et al.ICSE 2025 · 10 citations
