Online Cascade Learning for Efficient Inference over Streams
Lunyiu Nie, Zhimin Ding, Erdong Hu, Christopher M. Jermaine, Swarat Chaudhuri
Abstract
Large Language Models (LLMs) have a natural role in answering complex queries about data streams, but the high computational cost of LLM inference makes them infeasible in many such tasks. We propose online cascade learning, the first approach to address this challenge. The objective here is to learn a"cascade"of models, starting with lower-capacity models (such as logistic regression) and ending with a powerful LLM, along with a deferral policy that determines the model to be used on a given input. We formulate the task of learning cascades online as an imitation-learning problem, where smaller models are updated over time imitating the collected LLM demonstrations, and give a no-regret algorithm for the problem. Experimental results across four benchmarks show that our method parallels LLMs in accuracy while cutting down inference costs by as much as 90% with strong robustness against input distribution shifts, underscoring its efficacy and adaptability in stream processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 445e7848-13db-4701-8ddc-3ad380acd59fCited by top-tier papers9
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok et al.NeurIPS 2024 · 113 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- Gatekeeper: Improving Model Cascades Through Confidence TuningStephan Rabanser, Nathalie Rauschmayr, Achin Kulshrestha, Petra Poklukar et al.NeurIPS 2025 · 10 citations
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for ReasoningAntonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang et al.NeurIPS 2025 · 10 citations
- Transcending Cost-Quality Tradeoff in Agent Serving via Session-AwarenessYanyu Ren, Li Chen, Dan Li, Xizheng Wang et al.NeurIPS 2025 · 5 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Language Model Cascades: Token-Level Uncertainty And BeyondNeha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat et al.ICLR 2024 · 119 citations
Related papers
- Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP SystemsNeeraj Varshney, Chitta BaralEMNLP 2022 · 11 citations
- Cascaded Language Models for Cost-Effective Human-AI Decision-MakingClaudio Fanconi, Mihaela van der SchaarNeurIPS 2025 · 12 citations
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du et al.ICLR 2024 · 153 citations
- Near-Optimal Online Deployment and Routing for Streaming LLMsShaoang Li, Jian LiICLR 2026 · 3 citations
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 7 citations
