Towards Optimal Caching and Model Selection for Large Model Inference
Banghua Zhu, Ying Sheng, Lianmin Zheng, Clark W. Barrett, Michael I. Jordan, Jiantao Jiao
Abstract
Large Language Models (LLMs) and other large foundation models have achieved noteworthy success, but their size exacerbates existing resource consumption and latency challenges. In particular, the large-scale deployment of these models is hindered by the significant resource requirements during inference. In this paper, we study two approaches for mitigating these challenges: employing a cache to store previous queries and learning a model multiplexer to choose from an ensemble of models for query processing. Theoretically, we provide an optimal algorithm for jointly optimizing both approaches to reduce the inference cost in both offline and online tabular settings. By combining a caching algorithm, namely Greedy Dual Size with Frequency (GDSF) or Least Expected Cost (LEC), with a model multiplexer, we achieve optimal rates in both offline and online settings. Empirically, simulations show that the combination of our caching and model multiplexing algorithms greatly improves over the baselines, with up to 50× improvement over the baseline when the ratio between the maximum cost and minimum cost is 100. Experiments on real datasets show a 4.3× improvement in FLOPs over the baseline when the ratio for FLOPs is 10, and a 1.8× improvement in latency when the ratio for average latency is 1.85.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generative Caching for Structurally Similar Prompts and ResponsesSarthak Chakraborty, Suman Nath, Xuchao Zhang, Chetan Bansal et al.NeurIPS 2025 · 5 citations
- Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online AdaptationXutong Liu, Baran Atalar, Xiangxiang Dai, Jinhang Zuo et al.INFOCOM 2026 · 2 citations
- Parameterized Complexity of Caching in NetworksRobert Ganian, Fionn Mc Inerney, Dimitra TsigkariAAAI 2025
- Offline Learning for Combinatorial Multi-armed BanditsXutong Liu, Xiangxiang Dai, Jinhang Zuo, Siwei Wang et al.ICML 2025
Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
Related papers
- Online Context Caching for Distributed Large Language Models ServingBin Gao, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou et al.INFOCOM 2025 · 2 citations
- LLM Query Scheduling with Prefix Reuse and Latency ConstraintsGregory Dexter, Shao Tang, Ata Fatahi Baarzi, Qingquan Song et al.NeurIPS 2025 · 10 citations
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 64 citations
