Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
Sumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina, Philip Torr, Lewis Hammond, Christian Schröder de Witt
摘要
Recent capability increases in large language models (LLMs) open up applications in which groups of communicating generative AI agents solve joint tasks. This poses privacy and security challenges concerning the unauthorised sharing of information, or other unwanted forms of agent coordination. Modern steganographic techniques could render such dynamics hard to detect. In this paper, we comprehensively formalise the problem of secret collusion in systems of generative AI agents by drawing on relevant concepts from both AI and security literature. We study incentives for the use of steganography, and propose a variety of mitigation measures. Our investigations result in a model evaluation framework that systematically tests capabilities required for various forms of secret collusion. We provide extensive empirical results across a range of contemporary LLMs. While the steganographic capabilities of current models remain limited, GPT-4 displays a capability jump suggesting the need for continuous monitoring of steganographic frontier model capabilities. We conclude by laying out a comprehensive research program to mitigate future risks of collusion between generative AI models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- Large language models can learn and generalize steganographic chain-of-thought under process supervisionRobert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev 等NeurIPS 2025 · 被引用 29 次
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang 等NeurIPS 2025 · 被引用 26 次
- Early Signs of Steganographic Capabilities in Frontier LLMsArtur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann 等ICLR 2026 · 被引用 22 次
- Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative MontageJinwei Hu, Xinmiao Huang, Youcheng Sun, Yi Dong 等ACL 2026 · 被引用 11 次
它引用的顶会 Paper27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
相关 Paper
- TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking AgentDominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas 等EMNLP 2025
- Invisible Safety Threat: Malicious Finetuning for LLM via SteganographyGuangnian Wan, Xinyin Ma, Gongfan Fang, Xinchao WangICLR 2026 · 被引用 4 次
- AI Sandbagging: Language Models can Strategically Underperform on EvaluationsTeun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown 等ICLR 2025
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang 等ICLR 2024 · 被引用 441 次
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman 等ICML 2026 · 被引用 6 次
