SpareLLM: Automatically Selecting Task-Specific Minimum-Cost Large Language Models under Equivalence Constraint
Saehan Jo, Immanuel Trummer
Abstract
We introduce SpareLLM, Selecting P assable A nd R esource- E fficient LLM s, a novel LLM framework designed to minimize the inference costs (i.e., resource-efficient) of large-scale NLP tasks while ensuring sufficient result quality (i.e., passable). It enables users to specify an equivalence constraint in terms of the equivalence of outputs to those of the most powerful LLM. SpareLLM then generates results that deviate from the outputs of this LLM only with a probability below a user-defined threshold. SpareLLM employs a profiling phase that evaluates the performance of multiple LLMs to identify those that meet the user-defined equivalence level. It optimizes the tradeoff between profiling overheads and the anticipated cost savings resulting from profiling. Moreover, SpareLLM further reduces inference costs by strategically leveraging a mix of LLMs. Our experiments on five real-world datasets show that SpareLLM achieves significant cost savings, up to 8.6x, while generating equivalent outputs in 90% of cases compared to GPT-4-Turbo. Compared to recent LLM cascading baselines, SpareLLM demonstrates a superior tradeoff between cost and accuracy, accounting for 91.1% and 83.8% of the points on the Pareto curve for OpenAI and Llama models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4a09fd22-a8c0-4a08-8d14-736220d77daeCited by top-tier papers5
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li et al.SIGMOD 2026 · 24 citations
- Multi-Objective Agentic Rewrites for Unstructured Data ProcessingLindsey Linxi Wei, Shreya Shankar, Sepanta Zeighami, Yeounoh Chung et al.VLDB 2026 · 15 citations
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 7 citations
- SCOPE: Cost-Efficient Model Selection for Compound AI Systems under Quality ConstraintsYiqian Huang, Shiqi Zhang, Tianyuan Jin, Xiaokui XiaoKDD 2026
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
Related papers
- Cut Costs, Not Accuracy: LLM-Powered Data Processing with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranSIGMOD 2026 · 1 citation
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang et al.ICLR 2025
- Landmark-Guided Policy Optimization for Multi-Objective Language Model SelectionMarcio Monteiro, Weichen Li, Puyu Wang, Marius Kloft et al.ICML 2026
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk ParkICLR 2026 · 3 citations
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
