Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
Haozhen Zhang, Tao Feng, Jiaxuan You
Abstract
The rapid emergence of diverse large language models (LLMs) has spurred the development of LLM routers that assign user queries to the most suitable model. However, existing LLM routers typically perform a single-round, one-to-one mapping (i.e., assigning each query to a single model in isolation), which limits their capability to tackle complex tasks that demand the complementary strengths of multiple LLMs. In this paper, we present Router-R1, a reinforcement learning (RL)-based framework that formulates multi-LLM routing and aggregation as a sequential decision process. Router-R1 instantiates the router itself as a capable LLM, leveraging its reasoning ability to interleave "think" actions (internal deliberation) with "route" actions (dynamic model invocation), and integrates each response into its evolving context. To facilitate learning, we employ a lightweight rule-based reward comprising format rewards, final outcome rewards, and a novel cost reward for optimizing the balance between performance and cost, opening a pathway toward enhancing performance-cost trade-offs via RL. Router-R1 also conditions only on simple model descriptors such as pricing, latency, and example performance, enabling strong generalization to unseen model selection. Experiments on seven general and multi-hop QA benchmarks show that Router-R1 outperforms several strong baselines, achieving superior performance while maintaining robust generalization and cost management. ulab-uiuc/Router-R1 Hugging Face Collection * Work done as an intern at University of Illinois at Urbana-Champaign 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0db15f06-8451-4ea7-b7c0-7b15c12b128cCited by top-tier papers10
- Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred SkillsJustin Chih-Yao Chen, Sukwon Yun, Elias Stengel-Eskin, Tianlong Chen et al.ICML 2026 · 28 citations
- RouterArena: An Open Platform for Comprehensive Comparison of LLM RoutersYifan Lu, Rixin Liu, Jiayi Yuan, Xingqi Cui et al.ICLR 2026 · 21 citations
- MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled BenchmarksZixuan Ke, Yifei Ming, Austin Xu, Ryan Chin et al.ICML 2026 · 15 citations
- R2-Router: A New Paradigm for LLM Routing with ReasoningJiaqi Xue, Qian Lou, Jiarong Xing, Heng HuangICML 2026 · 12 citations
- GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMsTao Feng, Haozhen Zhang, Zijie Lei, Peixuan Han et al.ICLR 2026 · 11 citations
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
Related papers
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response TheoryWei Song, Zhenya Huang, Cheng Cheng, Weibo Gao et al.ACL 2025 · 20 citations
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang et al.ICLR 2025
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question AnsweringZheyuan Zhang, Kaiwen Shi, Zhengqing Yuan, Zehong Wang et al.ACL 2026
- Reward Model Routing in AlignmentXinle Wu, Yao LuICLR 2026 · 3 citations
- Lookahead Routing for Large Language ModelsCanbin Huang, Tianyuan Shi, Yuhua Zhu, Ruijun Chen et al.NeurIPS 2025 · 5 citations
