MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and Inference
Kaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng, Xuekai Zhu, Guoli Jia, Yuchen Fan, Xingtai Lv, Yuxin Zuo, Che Jiang, Yuru wang, Jianyu Wang
摘要
We present MARTI (Multi-Agent Reinforced Training and Inference), an opensource framework designed to facilitate scalable and efficient learning of multiagent LLM systems. MARTI supports centralized multi-agent interactions and distributed policy training, with the added capability of multi-turn asynchronous rollouts to enhance training efficiency. The framework includes dynamic workflows for multi-agent interactions, which integrate both rule-based verifiable rewards and LLM-based generative rewards. We validate the effectiveness of MARTI through comprehensive experiments on diverse mathematical tasks, demonstrating that multi-agent LLM-based systems outperform single-agent systems within the same inference budget after convergence. Our contributions lay the foundation for exploring scalable collaborations within LLM-based multi-agent systems and advancing the capabilities of large reasoning models. INTRODUCTION Large Reasoning Models (LRMs), such as DeepSeek-R1 (Guo et al., 2025) and OpenAI o1/o3 (El-Kishky et al., 2025) , highlight the significant role Reinforcement Learning (RL) plays in enhancing the reasoning capabilities of Large Language Models (LLMs) for solving complex problems. Notably, LRMs can explore and generate extended chains of thought using only rule-based outcome rewards. This RL paradigm has also demonstrated considerable progress in other domains, including visual reasoning (Liu et al., 2025d; Zhou et al., 2025; Team et al., 2025) and agentic reasoning (Wang et al., 2025c; Jin et al., 2025) tasks. These studies indicate the effectiveness of scaling up test-time inference computations using RL. However, further performance improvements through post-training RL typically demand substantial computational resources. Additionally, recent research suggests that RL primarily activates intrinsic capabilities and reflective patterns established during pre-training (Gandhi et al., 2025; Yue et al., 2025a; Shah et al., 2025) . Consequently, the initial model's passk performance sets an upper bound for RL-based enhancements (Yue et al., 2025a), which means the base model determines the reasoning limit. Therefore, the most viable approach for significantly boosting policy model performance remains within the scaling laws (Kaplan et al., 2020; Brown et al., 2020), either by training models on larger datasets or increasing the model's parameter size. Regarding the reinforcement learning stage, effectively leveraging the potential of exploration and environmental interaction remains a critical challenge (Silver & Sutton, 2025) . Meanwhile, LLM-based Multi-Agent Systems (MAS) (Han et al., 2024; Guo et al., 2024) scale inference computation by expanding the number of agents, each adaptively responding to specific tasks. Numerous open-source frameworks for LLM-based MAS are currently available, including AutoGen (Wu et al., 2023a), CAMEL (Li et al., 2023), and MetaGPT (Hong et al., 2024). However, these frameworks predominantly rely on LLM inference. This reliance makes their efficacy highly
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Decentralized LLM Collaboration with Multi-Agent Actor CriticShuo Liu, Tianle Chen, Ryan Amiri, Christopher AmatoICML 2026 · 被引用 6 次
- Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side VerificationHaoyang Liu, Jie Wang, Boxuan Niu, Xiongwei Han 等ICML 2026 · 被引用 4 次
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math ReasoningDan Qiao, Binbin Chen, Fengyu Cai, Jianlong Chen 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement LearningZiyu Wan, Yunxiang Li, Xiaoyu Wen, Yan Song 等NeurIPS 2025 · 被引用 76 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
- AgentPO: Enhancing Multi-Agent Collaboration via Reinforcement LearningLin Sun, Chuang Liu, Can Zhang, Yubin Wu 等ICLR 2026
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning EngineeringZexi Liu, Jingyi Chai, Xinyu Zhu, shuo tang 等ICML 2026
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to DeliberationZhiwei Zhang, Xiaomin Li, Yudi Lin, Hui Liu 等ICLR 2026 · 被引用 13 次
