Causal LLM Routing: End-to-End Regret Minimization from Observational Data
Asterios Tsiourvas, Wei Sun, Georgia Perakis
摘要
LLM routing aims to select the most appropriate model for each query, balancing competing performance metrics such as accuracy and cost across a pool of language models. Prior approaches typically adopt a decoupled strategy, where the metrics are first predicted and the model is then selected based on these estimates. This setup is prone to compounding errors and often relies on full-feedback data, where each query is evaluated by all candidate models, which is costly to obtain and maintain in practice. In contrast, we learn from observational data, which records only the outcome of the model actually deployed. We propose a causal end-to-end framework that learns routing policies by minimizing decision-making regret from observational data. To enable efficient optimization, we introduce two theoretically grounded surrogate objectives: a classification-based upper bound, and a softmax-weighted regret approximation shown to recover the optimal policy at convergence. We further extend our framework to handle heterogeneous cost preferences via an interval-conditioned architecture. Experiments on public benchmarks show that our method outperforms existing baselines, achieving state-of-the-art performance across different embedding models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DiSRouter: Distributed Self-Routing for LLM SelectionsHang Zheng, Hongshen Xu, Yongkai.lin, Shuai Fan 等ICLR 2026 · 被引用 6 次
- Scaling Small Agents Through Strategy AuctionsLisa Alazraki, Shen, Yoram Bachrach, Akhil MathurICML 2026 · 被引用 2 次
- Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert SelectionTianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Shi-Xiong Zhang 等ACL 2026 · 被引用 1 次
- CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLMSon Nguyen, Xinyuan Liu, Ransalu SenanayakeICML 2026
它引用的顶会 Paper5
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri 等ICLR 2024 · 被引用 172 次
- Fusing Models with Complementary ExpertiseHongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu 等ICLR 2024 · 被引用 44 次
- ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesKeke Huang, Yimin Shi, Dujian Ding, Yifei Li 等VLDB 2025 · 被引用 18 次
- Counterfactual Prediction for Outcome-Oriented TreatmentsHao Zou, Bo Li, Jiangang Han, Shuiping Chen 等ICML 2022 · 被引用 8 次
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang 等ICLR 2025
相关 Paper
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM RoutingYichi Zhang, Fangzheng Xie, Shu Yang, Chong WuICLR 2026 · 被引用 1 次
- Lookahead Routing for Large Language ModelsCanbin Huang, Tianyuan Shi, Yuhua Zhu, Ruijun Chen 等NeurIPS 2025 · 被引用 5 次
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang 等ICML 2025
- Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix OptimizationHaochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu 等ACL 2026 · 被引用 3 次
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
