Attack the Messages, Not the Agents: A Multi-round Adaptive Stealthy Tampering Framework for LLM-MAS
Bingyu Yan, Xiaoming Zhang, Ziyi Zhou, Chaozhuo Li, Ruilin Zeng, Yirui Qi, Tianbo Wang, Litian Zhang
摘要
Large language model-based multi-agent systems (LLM-MAS) effectively accomplish complex and dynamic tasks through inter-agent communication, but this reliance introduces substantial safety vulnerabilities. Existing attack methods targeting LLM-MAS either compromise agent internals or rely on direct and overt persuasion, which limit their effectiveness, adaptability, and stealthiness. In this paper, we propose MAST, a Multi-round Adaptive Stealthy Tampering framework designed to exploit communication vulnerabilities within the system. MAST integrates Monte Carlo Tree Search with Direct Preference Optimization to train an attack policy model that adaptively generates effective multi-round tampering strategies. Furthermore, to preserve stealthiness, we impose dual semantic and embedding similarity constraints during the tampering process. Comprehensive experiments across diverse tasks, communication architectures, and LLMs demonstrate that MAST consistently achieves high attack success rates while significantly enhancing stealthiness compared to baselines. These findings highlight the effectiveness, stealthiness, and adaptability of MAST, underscoring the need for robust communication safeguards in LLM-MAS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly DetectionJunjun Pan, Yixin Liu, Rui Miao, Kaize Ding 等ACL 2026 · 被引用 6 次
- Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration SystemsZiyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia TsvetkovACL 2026 · 被引用 1 次
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang 等ICML 2026
- Securing Multi-Agent Systems Against Corruptions via Node Contribution BackpropagationChengcan Wu, Zhixin Zhang, Mingqian Xu, Zeming Wei 等ICML 2026
- Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MASBingyu Yan, Xiaoming Zhang, Jinyu Hou, Chaozhuo Li 等ACL 2026
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue ResolutionWei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang 等NeurIPS 2024 · 被引用 210 次
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等NeurIPS 2024 · 被引用 195 次
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
相关 Paper
- ResMAS: Resilience Optimization in LLM-based Multi-agent SystemsZhilun Zhou, Zihan Liu, Jiahe Liu, Qingyu Shao 等AAAI 2026 · 被引用 2 次
- DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsXu Zhang, Xunjian Yin, Dinghao Jing, Huixuan Zhang 等EMNLP 2025 · 被引用 2 次
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 被引用 37 次
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM SystemsIshan Kavathekar, Hemang Jain, Ameya Rathod, Ponnurangam Kumaraguru 等ACL 2026 · 被引用 16 次
- Agents Under Siege: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt AttacksRana Muhammad Shahroz, Zhen Tan, Sukwon Yun, Charles Fleming 等ACL 2025 · 被引用 18 次
