A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns
Tianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen, Kang Liu, Jun Zhao
Abstract
With the development of large language models, they are widely used as agents in various fields. A key component of agents is memory, which stores vital information but is susceptible to jailbreak attacks. Existing research mainly focuses on single-agent attacks and shared memory attacks. However, real-world scenarios often involve independent memory. In this paper, we propose the Troublemaker Makes Chaos in Honest Town (TMCHT) task, a large-scale, multi-agent, multi-topology textbased attack evaluation framework. TMCHT involves one attacker agent attempting to mislead an entire society of agents. We identify two major challenges in multi-agent attacks: (1) Non-complete graph structure, (2) Largescale systems. We attribute these challenges to a phenomenon we term toxicity disappearing. To address these issues, we propose an Adversarial Replication Contagious Jailbreak (ARCJ) method, which optimizes the retrieval suffix to make poisoned samples more easily retrieved and optimizes the replication suffix to make poisoned samples have contagious ability. We demonstrate the superiority of our approach in TMCHT, with 23.51%, 18.95%, and 52.93% improvements in line, star topologies, and 100agent settings. It reveals potential contagion risks in widely used multi-agent architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54be45ba-e7dd-4a97-aad9-69a1b53bfae4Builds on7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du et al.ICML 2024 · 128 citations
Related papers
- Distract Large Language Models for Automatic Jailbreak AttackZeguan Xiao, Yan Yang, Guanhua Chen, Yun ChenEMNLP 2024 · 8 citations
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing et al.EMNLP 2024 · 6 citations
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu et al.USENIX Security 2025
- JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual SteeringRenmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan et al.ACM MM 2025 · 5 citations
- NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM JailbreaksJavad Rafiei Asl, Sidhant Narula, Mohammad GhasemiGol, Eduardo Blanco et al.EMNLP 2025 · 1 citation
