ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, Lei Bai
Abstract
Multi-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving. However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process. The core of ReSo is the proposed Collaborative Reward Model, which can provide fine-grained reward signals for MAS cooperation for optimization. We also introduce an automated data synthesis framework for generating MAS benchmarks, without human annotations. Experimentally, ReSo matches or outperforms existing methods. ReSo achieves 33.7% and 32.3% accuracy on Math-MAS and SciBench-MAS SciBench, while other methods completely fail. The code and data are available at Reso.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc383532-1008-4462-b51d-7da94de28790Cited by top-tier papers8
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- G-Memory: Tracing Hierarchical Memory for Multi-Agent SystemsGuibin Zhang, Muxin Fu, Kun Wang, Frank Wan et al.NeurIPS 2025 · 108 citations
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction RewardsXiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang et al.ICLR 2026 · 29 citations
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang et al.ACL 2026 · 17 citations
- Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual LearningWei Yang, Defu Cao, Jiacheng Pang, Muyan Weng et al.ICLR 2026 · 11 citations
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
Related papers
- Assemble Your Crew: Automatic Multi-agent Communication Topology Design via Autoregressive Graph GenerationShiyuan Li, Yixin Liu, Qingsong Wen, Chengqi Zhang et al.AAAI 2026 · 29 citations
- MASPO: Joint Prompt Optimization for LLM-based Multi-Agent SystemsZhexuan Wang, Xuebo Liu, Li Wang, Zifei Shan et al.ICML 2026 · 3 citations
- ResMAS: Resilience Optimization in LLM-based Multi-agent SystemsZhilun Zhou, Zihan Liu, Jiahe Liu, Qingyu Shao et al.AAAI 2026 · 2 citations
- Agentic Proposing: Enhancing Large language Model Reasoning via Compositional Skill SynthesisZhengbo Jiao, Shaobo Wang, Zifan Zhang, Xuan Ren et al.ICML 2026 · 10 citations
- HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent SystemsTianjun Yao, Zhaoyi Li, Zhiqiang ShenICML 2026 · 1 citation
