Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
Sumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip Torr, Lewis Hammond, Christian Schröder de Witt
Abstract
Recent capability increases in large language models (LLMs) open up applications in which groups of communicating generative AI agents solve joint tasks. This poses privacy and security challenges concerning the unauthorised sharing of information, or other unwanted forms of agent coordination. Modern steganographic techniques could render such dynamics hard to detect. In this paper, we comprehensively formalise the problem of secret collusion in systems of generative AI agents by drawing on relevant concepts from both AI and security literature. We study incentives for the use of steganography, and propose a variety of mitigation measures. Our investigations result in a model evaluation framework that systematically tests capabilities required for various forms of secret collusion. We provide extensive empirical results across a range of contemporary LLMs. While the steganographic capabilities of current models remain limited, GPT-4 displays a capability jump suggesting the need for continuous monitoring of steganographic frontier model capabilities. We conclude by laying out a comprehensive research program to mitigate future risks of collusion between generative AI models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e38b3a5f-926c-4c8e-9082-9421a3bb2106Cited by top-tier papers13
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes et al.NeurIPS 2025 · 52 citations
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev et al.NeurIPS 2025 · 29 citations
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
- Early Signs of Steganographic Capabilities in Frontier LLMsArtur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann et al.ICLR 2026 · 22 citations
- Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative MontageJinwei Hu, Xinmiao Huang, Youcheng Sun, Yi Dong et al.ACL 2026 · 11 citations
Builds on27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
Related papers
- TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking AgentDominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas et al.EMNLP 2025
- Invisible Safety Threat: Malicious Finetuning for LLM via SteganographyGuangnian Wan, Xinyin Ma, Gongfan Fang, Xinchao WangICLR 2026 · 4 citations
- AI Sandbagging: Language Models can Strategically Underperform on EvaluationsTeun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown et al.ICLR 2025
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman et al.ICML 2026 · 6 citations
