DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
Rui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin, Zixin Chen, Huamin Qu, Furui Cheng
摘要
Large language model (LLM)-based multi-agent systems have demonstrated impressive capabilities in handling complex tasks. However, the complexity of agentic behaviors makes these systems difficult to understand. When failures occur, developers often struggle to identify root causes and to determine actionable paths for improvement. Traditional methods that rely on inspecting raw log records are inefficient, given both the large volume and complexity of data. To address this challenge, we propose a framework and an interactive system, DiLLS, designed to reveal and structure the behaviors of multi-agent systems. The key idea is to organize information across three levels of query completion: activities, actions, and operations. By probing the multi-agent system through natural language, DiLLS derives and organizes information about planning and execution into a structured, multi-layered summary. Through a user study, we show that DiLLS significantly improves developers’ effectiveness and efficiency in identifying, diagnosing, and understanding failures in LLM-based multi-agent systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 等ICLR 2024 · 被引用 594 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
相关 Paper
- Why Do LLM-based Web Agents Fail? A Hierarchical Planning PerspectiveMohamed Aghzal, Gregory J. Stein, Ziyu YaoACL 2026 · 被引用 5 次
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang 等ICLR 2026 · 被引用 24 次
- Scope Delineation Before Localization: A Two-Stage Framework for Enhancing Failure Attribution in Multi-Agent SystemsKai Sun, Wenqiang Li, Bo Dong, Yuxin Lin 等AAAI 2026
- RCAFlow: A Workflow-Informed Hierarchical Planning Multi-Agent System for Root Cause AnalysisYufei Gao, Zhengong Cai, Bowei YangAAAI 2026
- Multi-Agent Design: Optimizing Agents with Better Prompts and TopologiesHan Zhou, Xingchen Wan, Ruoxi Sun, Hamid Palangi 等ICLR 2026 · 被引用 127 次
