CAM: A Causality-Based Analysis Framework for Multi-agent Code Generation Systems
Zongyi Lyu, Zhenlan Ji, Songqiang Chen, Liwen Wang, Yuheng Huang, Shuai Wang, Shing-Chi Cheung
摘要
Despite the remarkable success that Multi-Agent Code Generation Systems (MACGS) have achieved, the inherent complexity of multi-agent architectures produces substantial volumes of intermediate outputs. To date, the individual importance of these intermediate outputs to the system correctness remains opaque, which impedes targeted optimization of MACGS designs. To address this challenge, we propose CAM, the first C ausality-based A nalysis framework for M ACGS that systematically quantifies the contribution of different intermediate features to system correctness. By comprehensively categorizing intermediate outputs and systematically simulating realistic errors on intermediate features, we identify the important features for system correctness and aggregate their importance rankings, facilitating comprehensive analysis of MACGS. We instantiate CAM on representative MACGS across multiple backend LLMs and datasets and conduct extensive empirical analysis on the identified importance rankings. Our analysis reveals intriguing findings: first, we uncover context-dependent features—features whose importance emerges mainly through interactions with other features, revealing that quality assurance for MACGS should move beyond module-level validation to incorporate cross-feature consistency checks; second, we reveal that hybrid backend MACGS with different backend LLMs assigned according to their relative strength achieves up to 7.3% Pass@1 improvement, underscoring hybrid architectures as a promising direction for future MACGS design. We further demonstrate CAM’s practical utility through two applications: (1) failure repair, which achieves a 73.6% success rate by optimizing top-3 importance-ranked features and (2) feature pruning, that reduces up to 33.6% intermediate token consumption with negligible or sometimes positive performance impact by pruning low-importance features. Our work provides actionable insights for MACGS design and deployment, establishing causality analysis as a powerful approach for understanding and improving MACGS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 等ICLR 2024 · 被引用 594 次
相关 Paper
- Causality-Aided Evaluation and Explanation of Large Language Model-Based Code GenerationZhenlan Ji, Pingchuan Ma, Zongjie Li, Zhaoyu Wang 等ISSTA 2025 · 被引用 1 次
- Enhancing LLM-based Quantum Code Generation with Multi-Agent Optimization and Quantum Error CorrectionCharlie Campbell, Hao Mark Chen, Wayne Luk, Hongxiang FanDAC 2025 · 被引用 7 次
- OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent CollaborationShijun Li, Hilaf Hasson, Joydeep GhoshICML 2026
- MaCTG: Multi-Agent Collaborative Thought Graph for Automatic ProgrammingZixiao Zhao, Jing Sun, Zhe Hou, Zhiyuan Wei 等ICSE 2026
- TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated CodeJiangping Huang, Wenguang Ye, Weisong Sun, Jian Zhang 等ICSE 2026 · 被引用 1 次
