AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
Yulang Chen, Haoxuan Peng, Jinyan Liu, Zichen Wen, Dongrui Liu, Linfeng Zhang
Abstract
Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable capabilities in complex tasks. However, manually designing optimal communication topologies is labor-intensive, while automated expansion methods often result in bloated structures with redundant agents, leading to excessive token consumption. To address this problem, we introduce AgentSlimming, a plug-and-play compression framework for graph-structured multi-agent workflows. Motivated by pruning and quantization in neural networks, AgentSlimming compresses workflows by first estimating the importance score of each agent with a hybrid mechanism, and then removes redundant agents or replaces them with low-cost ones, where each operation is validated using a baseline-anchored acceptance rule to prevent performance collapse. Experiments show that AgentSlimming reduces average token cost by up to 78.9% with negligible performance degradation, and sometimes even improves accuracy, achieving a strong Pareto-optimal trade-off between cost and quality. Our code is publicly available at https://github.com/CitrusYL/AgentSlimming
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7abc3410-e541-4ade-a008-028733a63ebdBuilds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent SystemsGuibin Zhang, Yanwei Yue, Zhixun Li, Sukwon Yun et al.ICLR 2025
- AgentTailor: A Semantic-Aware LLM-Based Multi-Agent System with Actor-Critic StructurePeiting Yang, Jiahao Shi, Caiyi Xu, Ming Liu et al.ICML 2026
- AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent CollaborationZhexuan Wang, Yutong Wang, Xuebo Liu, Liang Ding et al.ACL 2025 · 32 citations
- ACON: Optimizing Context Compression for Long-horizon LLM AgentsMinki Kang, Wei-Ning Chen, Dongge Han, Huseyin Inan et al.ICML 2026
- LLM-as-Scheduler: Agentic Workflow Dynamic SchedulingDawei Xiang, Kexin Chu, Wenyan Xu, Wenhui Zhang et al.ACL 2026
