Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, Min Lin
摘要
A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further intervention from the adversary, (almost) all agents will become infected exponentially fast and exhibit harmful behaviors. To validate the feasibility of infectious jailbreak, we simulate multi-agent environments containing up to one million LLaVA-1.5 agents, and employ randomized pair-wise chat as a proof-of-concept instantiation for multi-agent interaction. Our results show that feeding an (infectious) adversarial image into the memory of any randomly chosen agent is sufficient to achieve infectious jailbreak. Finally, we derive a simple principle for determining whether a defense mechanism can provably restrain the spread of infectious jailbreak, but how to design a practical defense that meets this principle remains an open question to investigate. Our code is available at https://github.com/sail-sg/Agent-Smith . * Equal contribution (ordered by dice rolling). The project was led by Tianyu Pang, and done during Xiangming Gu and Xiaosen Zheng's internships at Sea AI Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language ModelsChristian Schlarmann, Naman Deep Singh, Francesco Croce, Matthias HeinICML 2024 · 被引用 114 次
- Best-of-N JailbreakingJohn Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer 等NeurIPS 2025 · 被引用 78 次
- SAGA: A Security Architecture for Governing AI Agentic SystemsGeorgios Syros, Anshuman Suri, Jacob Ginesin, Cristina Nita-Rotaru 等NDSS 2026 · 被引用 63 次
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang 等NeurIPS 2025 · 被引用 49 次
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan 等ACL 2025 · 被引用 37 次
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan 等CVPR 2025
- A Troublemaker with Contagious Jailbreak Makes Chaos in Honest TownsTianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen 等ACL 2025
- DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsXu Zhang, Xunjian Yin, Dinghao Jing, Huixuan Zhang 等EMNLP 2025 · 被引用 2 次
- PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of mUlti-turn jailbrEaksNeeladri Bhuiya, Madhav Aggarwal, Diptanshu PurwarICLR 2026 · 被引用 2 次
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu 等ACL 2025 · 被引用 14 次
