SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality Constraints
Yiqian Huang, Shiqi Zhang, Tianyuan Jin, Xiaokui Xiao
摘要
A compound AI system consists of multiple LLM modules, together handling complex and multi-step tasks that exceed the capabilities of a single model. Existing systems often use a single expensive LLM across all modules to improve the result quality of the whole system. However, this configuration incurs prohibitive costs, particularly for data management and analytics tasks at scale, such as data manipulation. To this end, we formalize the problem of constrained LLM selection for compound AI systems, leveraging the diverse pricing and capabilities of different LLMs to achieve competitive quality at lower cost. Given a query dataset and a user-specified quality threshold, we aim to select an LLM for each module to minimize the system's average cost while ensuring that overall quality meets the required threshold. To solve this problem, we propose SCOPE, a cost-efficient optimization algorithm. Unlike existing approaches that rely on expensive dataset-level evaluations, SCOPE exploits per-query results to rapidly estimate the system's cost and quality, and constructs confidence bounds to guide the search for promising LLM combinations. Furthermore, SCOPE provides theoretical guarantees for meeting the quality threshold and achieving near-optimal average cost. We evaluate SCOPE against 7 baselines on three data processing tasks, demonstrating that it outperforms all baselines. Under the same search budget and quality constraint, it finds solutions with up to 20× lower cost than the best competitor during the search and achieves up to 6× lower final cost in the returned solution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Constrained Efficient Global Optimization of Expensive Black-box FunctionsWenjie Xu, Yuning Jiang, Bratislav Svetozarevic, Colin N. JonesICML 2023 · 被引用 1,916 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang 等ICLR 2024 · 被引用 170 次
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du 等ICLR 2024 · 被引用 153 次
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 被引用 143 次
相关 Paper
- Conformal Constrained Policy Optimization for Cost-Effective LLM AgentsWenwen Si, Sooyong Jang, Insup Lee, Osbert BastaniAAAI 2026 · 被引用 3 次
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc ReasoningQi Cao, Shuhao Zhang, Ruizhe Zhou, Ruiyi Zhang 等ICML 2026 · 被引用 2 次
- Optimas: Optimizing Compound AI Systems with Globally Aligned Local RewardsShirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee 等ICLR 2026 · 被引用 25 次
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano 等VLDB 2026 · 被引用 9 次
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang 等VLDB 2026 · 被引用 5 次
