CONCUR: A Framework for Continual Constrained and Unconstrained Routing
Peter Baile Chen, Weiyue Li, Dan Roth, Mike Cafarella, Samuel Madden, Jacob Andreas
Abstract
AI tasks differ in complexity and are best addressed with different computation strategies (e.g., combinations of models and decoding methods). Hence, an effective routing system that maps tasks to the appropriate strategies is crucial. Most prior methods build the routing framework by training a single model across all strategies, which demands full retraining whenever new strategies appear and leads to high overhead. Attempts at such continual routing, however, often face difficulties with generalization. Prior models also typically use a single input representation, limiting their ability to capture the full complexity of the routing problem and leading to sub-optimal routing decisions. To address these gaps, we propose CONCUR, a continual routing framework that supports both constrained and unconstrained routing (i.e., routing with or without a budget). Our modular design trains a separate predictor model for each strategy, enabling seamless incorporation of new strategies with low additional training cost. Our predictors also leverage multiple representations of both tasks and computation strategies to better capture overall problem complexity. Experiments on both in-distribution and out-of-distribution, knowledge- and reasoning-intensive tasks show that our method outperforms the best single strategy and strong existing routing techniques with higher end-to-end accuracy and lower inference cost in both continual and non-continual settings, while also reducing training cost in the continual setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok et al.NeurIPS 2024 · 113 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
Related papers
- R2-Router: A New Paradigm for LLM Routing with ReasoningJiaqi Xue, Qian Lou, Jiarong Xing, Heng HuangICML 2026 · 12 citations
- Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction TuningKunlun Xu, YanQin Zhang, Wenwen Qiang, Jiahuan ZhouICML 2026
- Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelsKai Tong, Kang Pan, Xiao Zhang, Erli Meng et al.ICCV 2025 · 3 citations
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc ReasoningQi Cao, Shuhao Zhang, Ruizhe Zhou, Ruiyi Zhang et al.ICML 2026 · 2 citations
- One Person, One Model, One World: Learning Continual User Representation without ForgettingFajie Yuan, Guoxiao Zhang, Alexandros Karatzoglou, Joemon M. Jose et al.SIGIR 2021 · 52 citations
