Interactive Debugging and Steering of Multi-Agent AI Systems
Will Epperson, Gagan Bansal, Victor C. Dibia, Adam Fourney, Jack Gerrits, Erkang (Eric) Zhu, Saleema Amershi
摘要
Fully autonomous teams of LLM-powered AI agents are emerging that collaborate to perform complex tasks for users. What challenges do developers face when trying to build and debug these AI agent teams? In formative interviews with five AI agent developers, we identify core challenges: difficulty reviewing long agent conversations to localize errors, lack of support in current tools for interactive debugging, and the need for tool support to iterate on agent configuration. Based on these needs, we developed an interactive multi-agent debugging tool, AGDebugger, with a UI for browsing and sending messages, the ability to edit and reset prior agent messages, and an overview visualization for navigating complex message histories. In a two-part user study with 14 participants, we identify common user strategies for steering agents and highlight the importance of interactive message resets for debugging. Our studies deepen understanding of interfaces for debugging increasingly important agentic workflows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous EnvironmentsRomain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja 等ICLR 2026 · 被引用 29 次
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent SystemsMing Ma, Jue Zhang, Fangkai Yang, Yu Kang 等ICLR 2026 · 被引用 24 次
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systemsYifan Yu, Moyan Li, Shaoyuan Xu, Jinmiao Fu 等ICML 2026 · 被引用 9 次
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 被引用 3 次
- Code with Me or for Me? How Increasing AI Automation Transforms Developer WorkflowsValerie Chen, Ameet Talwalkar, Robert Brennan, Graham NeubigCHI 2026 · 被引用 2 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
相关 Paper
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen 等ACL 2024
- An LLM-based multi-agent framework for agile effort estimationThanh-Long Bui, Hoa Khanh Dam, Rashina HodaASE 2025 · 被引用 2 次
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin 等CHI 2026 · 被引用 2 次
- ChatDBG: Augmenting Debugging with Large Language ModelsKyla Levin, Nicolas van Kempen, Emery D. Berger, Stephen N. FreundFSE 2025 · 被引用 3 次
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
