Causal LLM Routing: End-to-End Regret Minimization from Observational Data
Asterios Tsiourvas, Wei Sun, Georgia Perakis
Abstract
LLM routing aims to select the most appropriate model for each query, balancing competing performance metrics such as accuracy and cost across a pool of language models. Prior approaches typically adopt a decoupled strategy, where the metrics are first predicted and the model is then selected based on these estimates. This setup is prone to compounding errors and often relies on full-feedback data, where each query is evaluated by all candidate models, which is costly to obtain and maintain in practice. In contrast, we learn from observational data, which records only the outcome of the model actually deployed. We propose a causal end-to-end framework that learns routing policies by minimizing decision-making regret from observational data. To enable efficient optimization, we introduce two theoretically grounded surrogate objectives: a classification-based upper bound, and a softmax-weighted regret approximation shown to recover the optimal policy at convergence. We further extend our framework to handle heterogeneous cost preferences via an interval-conditioned architecture. Experiments on public benchmarks show that our method outperforms existing baselines, achieving state-of-the-art performance across different embedding models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff5af9c7-71c0-4400-ae10-b79cfb14b4a5Cited by top-tier papers4
- DiSRouter: Distributed Self-Routing for LLM SelectionsHang Zheng, Hongshen Xu, Yongkai.lin, Shuai Fan et al.ICLR 2026 · 6 citations
- Scaling Small Agents Through Strategy AuctionsLisa Alazraki, Shen, Yoram Bachrach, Akhil MathurICML 2026 · 2 citations
- Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert SelectionTianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Shi-Xiong Zhang et al.ACL 2026 · 1 citation
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
Builds on5
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri et al.ICLR 2024 · 172 citations
- Fusing Models with Complementary ExpertiseHongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu et al.ICLR 2024 · 44 citations
- ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesKeke Huang, Yimin Shi, Dujian Ding, Yifei Li et al.VLDB 2025 · 18 citations
- Counterfactual Prediction for Outcome-Oriented TreatmentsHao Zou, Bo Li, Jiangang Han, Shuiping Chen et al.ICML 2022 · 8 citations
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang et al.ICLR 2025
Related papers
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM RoutingYichi Zhang, Fangzheng Xie, Shu Yang, Chong WuICLR 2026 · 1 citation
- Lookahead Routing for Large Language ModelsCanbin Huang, Tianyuan Shi, Yuhua Zhu, Ruijun Chen et al.NeurIPS 2025 · 5 citations
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang et al.ICML 2025
- Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix OptimizationHaochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu et al.ACL 2026 · 3 citations
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
