Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
Junjun Pan, Yixin Liu, Rui Miao, Kaize Ding, Yu Zheng, Quoc Viet Hung Nguyen, Alan Wee-Chung Liew, Shirui Pan
Abstract
Large language model (LLM)-based multi-agent systems (MAS) have shown strong capabilities in solving complex tasks. As MAS become increasingly autonomous in various safety-critical tasks, detecting malicious agents has become a critical security concern. Although existing graph anomaly detection (GAD)-based defenses can identify anomalous agents, they mainly rely on coarse sentence-level information and overlook fine-grained lexical cues, leading to suboptimal performance. Moreover, the lack of interpretability in these methods limits their reliability and real-world applicability. To address these limitations, we propose XG-Guard, an explainable and fine-grained safeguarding framework for detecting malicious agents in MAS. To incorporate both coarse and fine-grained textual information for anomalous agent identification, we utilize a bi-level agent encoder to jointly model the sentence- and token-level representations of each agent. A theme-based anomaly detector further captures the evolving discussion focus in MAS dialogues, while a bi-level score fusion mechanism quantifies token-level contributions for explanation. Extensive experiments across diverse MAS topologies and attack scenarios demonstrate robust detection performance and strong interpretability of XG-Guard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9898aa1e-d3b4-4938-91f6-6d8ac4d43b8fCited by top-tier papers3
- OFA-MAS: One-for-All Multi-Agent System Topology Design based on Mixture-of-Experts Graph Generative ModelsShiyuan Li, Yixin Liu, Yu Zheng, Mei Li et al.WWW 2026 · 2 citations
- Rethinking Feature Alignment in Generalist Graph Anomaly Detection: A Relational Fingerprint-based ApproachYujing Liu, Yixin Liu, Yu Zheng, Alan Liew et al.ICML 2026 · 1 citation
- Escaping the Homophily Trap: A Threshold-free Graph Outlier Detection Framework via Clustering-guided Edge ReweightingYunhe Zhang, Jinyu Cai, Qi Hao, Pengyang Wang et al.ICLR 2026
Builds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- A-Mem: Agentic Memory for LLM AgentsWujiang Xu, Zujie Liang, Kai Mei, Hang Gao et al.NeurIPS 2025 · 1,138 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- Truncated Affinity Maximization: One-class Homophily Modeling for Graph Anomaly DetectionHezhe Qiao, Guansong PangNeurIPS 2023 · 84 citations
Related papers
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent SystemsShilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan et al.ACL 2025 · 37 citations
- GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph ModelingJialong Zhou, Lichao Wang, Xiao YangNeurIPS 2025 · 40 citations
- GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language ModelYunhe Pang, Bo Chen, Fanjin Zhang, Yanghui Rao et al.KDD 2025
- Securing Multi-Agent Systems Against Corruptions via Node Contribution BackpropagationChengcan Wu, Zhixin Zhang, Mingqian Xu, Zeming Wei et al.ICML 2026
- When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent SystemsHaowen Xu, Xue Tan, Lei Ma, Zhihao Zhang et al.ICML 2026
