OctoSelector: Efficient and Effective Batch-Aware Model Selection for Large Language Models
Guangxue Zhang, Yiming Lin, Sharad Mehrotra
摘要
Large Language Models (LLMs) vary significantly in metrics such as accuracy, latency, and cost, making it challenging for users and applications to decide which model to invoke for each query. This paper presents O cto S elector , a framework for LLM selection that satisfies user-defined objectives and constraints across multiple metrics. In the pre-processing phase, O cto S elector learns difficulty-aware representations of queries based on both input and output complexity, clustering them into similar difficulty groups to enable efficient performance estimation across multiple LLMs. During inference, O cto S elector supports LLM selection for batched workload, formulating it as an Integer Linear Programming (ILP) problem that optimizes a user-defined objective (e.g., minimizing cost or latency, or maximizing accuracy) while enforcing constraints on other metrics. We evaluate O cto S elector on two types of tasks: NL2SQL using the Spider and BIRD benchmarks, and sentiment analysis using the IMDb benchmark. When optimizing for cost under accuracy and latency constraints, O cto S elector achieves up to a 67.7% cost reduction on NL2SQL tasks for batched workloads compared to state-of-the-art approaches.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PRISM: Navigating Cost-Accuracy Trade-offs for NL2SQLGaurav Tarlok Kakkar, Yeounoh Chung, Fatma Özcan, Stephen Mussmann 等SIGMOD 2026
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response TheoryWei Song, Zhenya Huang, Cheng Cheng, Weibo Gao 等ACL 2025 · 被引用 20 次
- MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level GuaranteesHerbert Woisetschläger, Ryan Zhang, Shiqiang Wang, Hans-Arno JacobsenNeurIPS 2025 · 被引用 11 次
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
- SELA: Smart Edge LLM Agent to Optimize Response Trade-offs of AI AssistantsShreshth Tuli, Giuliano Casale, Manuel RoveriUbiComp 2025 · 被引用 5 次
