RouteLLM: Learning to Route LLMs from Preference Data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, Ion Stoica
Abstract
Large language models (LLMs) excel at a wide range of tasks, but choosing the right model often involves balancing performance and cost. Powerful models offer better results but are expensive, while smaller models are more cost-effective but less capable. To address this trade-off, we introduce a training framework for learning efficient router models that dynamically select between a stronger and weaker LLM during inference. Our framework leverages human preference data and employs data augmentation techniques to enhance performance. Evaluations on public benchmarks show that our approach can reduce costs by over 2 times without sacrificing response quality. Moreover, our routers exhibit strong generalization capabilities, maintaining performance even when routing between LLMs not included in training. This highlights the potential of our framework to deliver cost-effective, high-performance LLM solutions. INTRODUCTION Recent advances in large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language tasks. From open-ended conversation and question answering to text summarization and code generation, LLMs have demonstrated an impressive level of fluency and understanding (Achiam et al., 2023; Bubeck et al., 2023) . This rapid progress has been enabled by a combination of architectural innovations, such as the Transformer architecture (Vaswani et al., 2017) , as well as scaling up data and training infrastructure (Brown et al., 2020; Radford et al., 2019) . However, not all LLMs are created equal-there exists wide variation in the sizes of different LLMs, which in turn affects the resources required to serve them. LLMs also differ in terms of the data on which they are trained, which in turn leads to variations in the strengths, weaknesses, and capabilities of different models. Broadly speaking, larger models tend to be more capable but come at a higher cost, while smaller models tend to be less capable but cheaper to serve. This heterogeneous landscape presents a dilemma in the practical deployment of LLMs. Although routing all user queries to the largest and most capable model ensures high-quality results, it is prohibitively expensive. Conversely, routing queries to smaller models can save costs-by more than 50x (e.g., Claude-3 Haiku vs. Opus 1 )-but may result in lower quality responses, as the smaller model may not handle complex queries effectively. LLM routing (Ding et al., 2024; Hu et al., 2024) offers an effective solution by first processing each user query through a router, which then determines the most suitable LLM to handle the query. The router can direct simpler queries to smaller models and more complex ones to larger models, thereby balancing response quality with cost efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 403bcc85-0693-4305-9ed7-af322f99bef5Cited by top-tier papers39
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement LearningHaozhen Zhang, Tao Feng, Jiaxuan YouNeurIPS 2025 · 81 citations
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token RoutingTianyu Fu, Yi Ge, Yichen You, Enshu Liu et al.NeurIPS 2025 · 32 citations
- Causal LLM Routing: End-to-End Regret Minimization from Observational DataAsterios Tsiourvas, Wei Sun, Georgia PerakisNeurIPS 2025 · 27 citations
- Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time ComputeJianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi et al.AAAI 2026 · 18 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
Related papers
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang et al.ICML 2025
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response TheoryWei Song, Zhenya Huang, Cheng Cheng, Weibo Gao et al.ACL 2025 · 20 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
- GraphRouter: A Graph-based Router for LLM SelectionsTao Feng, Yanzhen Shen, Jiaxuan YouICLR 2025
