MARS²: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation
Pengfei Li, Shijie Wang, Fangyuan Li, Yikun Fu, Kaifeng Liu, Kaiyan Zhang, Dazhi Zhang, Yuqiang Li, Biqing Qi, Bowen Zhou
Abstract
Reinforcement learning (RL) paradigms have demonstrated strong performance on reasoning-intensive tasks such as code generation. However, limited trajectory diversity often leads to diminishing returns, which constrains the achievable performance ceiling. Search-enhanced RL alleviates this issue by introducing structured exploration, which remains constrained by the single-agent policy priors. Meanwhile, leveraging multiple interacting policies can acquire more diverse exploratory signals, but existing approaches are typically decoupled from structured search. We propose MARS (Multi-Agent Reinforced Tree-Search Scaling), a unified RL framework in which multiple independently-optimized agents collaborate within a shared tree-structured search environment. MARS models the search tree as a learnable multi-agent interaction environment, enabling heterogeneous agents to collaboratively generate and refine candidate solutions within a shared search topology. To support effective learning, we introduce a path-level group advantage formulation based on tree-consistent reward shaping, which facilitates effective credit assignment across complex search trajectories. Experiments on code generation benchmarks show that MARS consistently improves performance across diverse model combinations and training settings, demonstrating the effectiveness of coupling multi-agent collaboration with tree search for enhancing reinforcement learning. Our code is publicly available at https://github.com/TsinghuaC3I/MARTI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f16aaee-3c4d-4c5e-b030-da2da34b8fcaBuilds on6
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree SearchYuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki et al.NeurIPS 2025 · 67 citations
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement LearningChanwoo Park, Seungju Han, Xingzhi Guo, Asuman E. Ozdaglar et al.ACL 2025 · 63 citations
- Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningJaehun Jung, Seungju Han, Ximing Lu, Skyler Hallinan et al.NeurIPS 2025 · 50 citations
Related papers
- AT²PO: Agentic Turn-based Policy Optimization via Tree SearchZefang Zong, Dingwei Chen, Yang Li, Qi Yi et al.ACL 2026 · 3 citations
- Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner CollaborationBowei He, Minda Hu, Zenan Xu, Hongru WANG et al.ICML 2026 · 9 citations
- TreeRL: LLM Reinforcement Learning with On-Policy Tree SearchZhenyu Hou, Ziniu Hu, Yujiang Li, Rui Lu et al.ACL 2025
- MARTI: A Framework for Multi-Agent LLM Systems Reinforced Training and InferenceKaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng et al.ICLR 2026
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMsYuanhao Li, Mingshan Liu, Hongbo Wang, Yiding Zhang et al.AAAI 2026
