Routing, Cascades, and User Choice for LLMs
Rafid Mahmood
Abstract
To mitigate the trade-offs between performance and costs, LLM providers route user tasks to different models based on task difficulty and latency. We study the effect of LLM routing with respect to user behavior. We propose a game between an LLM provider with two models (standard and reasoning) and a user who can re-prompt or abandon tasks if the routed model cannot solve them. The user's goal is to maximize their utility minus the delay from using the model, while the provider minimizes the cost of servicing the user. We solve this Stackelberg game by fully characterizing the user best response and simplifying the provider problem. We observe that in nearly all cases, the optimal routing policy involves a static policy with no cascading that depends on the expected utility of the models to the user. Furthermore, we reveal a misalignment gap between the provider-optimal and user-preferred routes when the user's and provider's rankings of the models with respect to utility and cost differ. Finally, we demonstrate conditions for extreme misalignment where providers are incentivized to throttle the latency of the models to minimize their costs, consequently depressing user utility. The results yield simple threshold rules for single-provider, single-user interactions and clarify when routing, cascading, and throttling help or harm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afd6b401-523e-4dd2-a5fa-df5f986d7963Builds on12
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- Low-Budget Active Learning via Wasserstein Distance: An Integer Programming ApproachRafid Mahmood, Sanja Fidler, Marc T. LawICLR 2022 · 44 citations
- Cost-of-Pass: An Economic Framework for Evaluating Language ModelsMehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yüksekgönül et al.ICLR 2026 · 40 citations
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
Related papers
- Pricing Online LLM Services with Data-Calibrated Stackelberg Routing GameZhendong Guo, Wenchao Bai, Jiahui JinAAAI 2026
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning TasksVansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal et al.ICLR 2026 · 7 citations
- Causal LLM Routing: End-to-End Regret Minimization from Observational DataAsterios Tsiourvas, Wei Sun, Georgia PerakisNeurIPS 2025 · 27 citations
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang et al.ICML 2025
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
