A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
Zihan Xu, Haolin Tian, Hai Jiang
摘要
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to improve inference-time efficiency. From the perspective of inference-time execution, this paper models parallelism in multi-agent systems as two distinct levels of decision processes: Replica Parallelism, which explores multiple complete solution paths at the task level, and Structural Parallelism, which enables concurrent execution within a single solution path through task decomposition. However, the roles of different forms of parallelism and their interrelationships still lack systematic study in terms of unified organization and coordination. We therefore propose TIPEX, a controllable execution framework that unifies these two levels of parallelism and coordinates their roles within the inference process under a unified execution semantics while supporting systematic combinations and analyses of different parallel strategies and parameter configurations. Systematic experiments on the GAIA benchmark demonstrate that inference-time parallelism can significantly improve accuracy and reduce end-to-end latency at the cost of increased token consumption. Further analysis shows that Replica and Structural Parallelism exhibit complementary effects across task complexities, with tasks of intermediate difficulty benefiting most from their coordination, while overly aggressive parallel strategies do not necessarily yield better performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsWeize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 等ICLR 2024 · 被引用 594 次
- An LLM Compiler for Parallel Function CallingSehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee 等ICML 2024 · 被引用 142 次
相关 Paper
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel ExecutionTianrui Qin, Qianben Chen, Sinuo Wang, He Xing 等ICLR 2026 · 被引用 28 次
- Parallelizing LLM Agent Execution with Contrastive Task AllocationYuyang Peng, Yanling Xu, Shuyi Wang, Xiaofei Liao 等KDD 2026 · 被引用 1 次
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language ModelsLong (Tony) Lian, Sida Wang, Felix Juefei-Xu, Tsu-Jui Fu 等ICML 2026
- Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool InvocationDongsheng Zhu, Weixian Shi, Zhengliang Shi, Zhaochun Ren 等ACL 2025 · 被引用 16 次
- Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative ExplorationShuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo 等OSDI 2026
