SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality Constraints
Yiqian Huang, Shiqi Zhang, Tianyuan Jin, Xiaokui Xiao
Abstract
A compound AI system consists of multiple LLM modules, together handling complex and multi-step tasks that exceed the capabilities of a single model. Existing systems often use a single expensive LLM across all modules to improve the result quality of the whole system. However, this configuration incurs prohibitive costs, particularly for data management and analytics tasks at scale, such as data manipulation. To this end, we formalize the problem of constrained LLM selection for compound AI systems, leveraging the diverse pricing and capabilities of different LLMs to achieve competitive quality at lower cost. Given a query dataset and a user-specified quality threshold, we aim to select an LLM for each module to minimize the system's average cost while ensuring that overall quality meets the required threshold. To solve this problem, we propose SCOPE, a cost-efficient optimization algorithm. Unlike existing approaches that rely on expensive dataset-level evaluations, SCOPE exploits per-query results to rapidly estimate the system's cost and quality, and constructs confidence bounds to guide the search for promising LLM combinations. Furthermore, SCOPE provides theoretical guarantees for meeting the quality threshold and achieving near-optimal average cost. We evaluate SCOPE against 7 baselines on three data processing tasks, demonstrating that it outperforms all baselines. Under the same search budget and quality constraint, it finds solutions with up to 20× lower cost than the best competitor during the search and achieves up to 6× lower final cost in the returned solution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a79b306-58a7-4867-b7aa-719a19dd6ec8Builds on16
- Constrained Efficient Global Optimization of Expensive Black-box FunctionsWenjie Xu, Yuning Jiang, Bratislav Svetozarevic, Colin N. JonesICML 2023 · 1,916 citations
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 909 citations
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang et al.ICLR 2024 · 170 citations
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du et al.ICLR 2024 · 153 citations
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
Related papers
- Conformal Constrained Policy Optimization for Cost-Effective LLM AgentsWenwen Si, Sooyong Jang, Insup Lee, Osbert BastaniAAAI 2026 · 3 citations
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc ReasoningQi Cao, Shuhao Zhang, Ruizhe Zhou, Ruiyi Zhang et al.ICML 2026 · 2 citations
- Optimas: Optimizing Compound AI Systems with Globally Aligned Local RewardsShirley Wu, Parth Sarthi, Shiyu Zhao, Aaron Lee et al.ICLR 2026 · 25 citations
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano et al.VLDB 2026 · 9 citations
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang et al.VLDB 2026 · 5 citations
