BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown Attacks
Rui Miao, Yixin Liu, Yili Wang, Xu Shen, Yue Tan, Yiwei Dai, Shirui Pan, Xin Wang
Abstract
The security of LLM-based multi-agent systems (MAS) is critically threatened by propagation vulnerability, where malicious agents can distort collective decision-making through inter-agent message interactions. While existing supervised defense methods demonstrate promising performance, they may be impractical in real-world scenarios due to their heavy reliance on labeled malicious agents to train a supervised malicious detection model. To enable practical and generalizable MAS defenses, in this paper, we propose BlindGuard, an unsupervised defense method that learns without requiring any attack-specific labels or prior knowledge of malicious behaviors. To this end, we establish a hierarchical agent encoder to capture individual, neighborhood, and global interaction patterns of each agent, providing a comprehensive understanding for malicious agent detection. Meanwhile, we design a corruption-guided detector that consists of directional noise injection and contrastive learning, allowing effective detection model training solely on normal agent behaviors. Extensive experiments show that BlindGuard effectively detects diverse attack types (i.e., prompt injection, memory poisoning, and tool attack) across MAS with various communication patterns while maintaining superior generalizability compared to supervised baselines. The code is available at: https://github.com/MR9812/BlindGuard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86fcc6c6-f1b7-4b04-918b-36f89ef37fdaCited by top-tier papers13
- Assemble Your Crew: Automatic Multi-agent Communication Topology Design via Autoregressive Graph GenerationShiyuan Li, Yixin Liu, Qingsong Wen, Chengqi Zhang et al.AAAI 2026 · 29 citations
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao et al.ICLR 2026 · 28 citations
- Where Graph Meets Heterogeneity: Multi-View Collaborative Graph ExpertsZhihao Wu, Jinyu Cai, Yunhe Zhang, Jielong Lu et al.NeurIPS 2025 · 6 citations
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly DetectionJunjun Pan, Yixin Liu, Rui Miao, Kaize Ding et al.ACL 2026 · 6 citations
- Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent SystemsXu Shen, Yixin Liu, Yiwei Dai, Yili Wang et al.EMNLP 2025 · 4 citations
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical ProblemsBin Lei, Yi Zhang, Shan Zuo, Ali Payani et al.NeurIPS 2024 · 57 citations
Related papers
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang et al.ICML 2026
- Securing Multi-Agent Systems Against Corruptions via Node Contribution BackpropagationChengcan Wu, Zhixin Zhang, Mingqian Xu, Zeming Wei et al.ICML 2026
- When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent SystemsLingxi Zhang, Guangtao Zheng, Hanjie ChenICML 2026 · 1 citation
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
- SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat DetectionYang Feng, Xudong PanWWW 2026 · 3 citations
