Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP Systems
Neeraj Varshney, Chitta Baral
摘要
Do all instances need inference through the big models for a correct prediction? Perhaps not; some instances are easy and can be answered correctly by even small capacity models. This provides opportunities for improving the computational efficiency of systems. In this work, we present an explorative study on 'model cascading', a simple technique that utilizes a collection of models of varying capacities to accurately yet efficiently output predictions. Through comprehensive experiments in multiple task settings that differ in the number of models available for cascading (K value), we show that cascading improves both the computational efficiency and the prediction accuracy. For instance, in K=3 setting, cascading saves up to 88.93% computation cost and consistently achieves superior prediction accuracy with an improvement of up to 2.18%. We also study the impact of introducing additional models in the cascade and show that it further increases the efficiency improvements. Finally, we hope that our work will facilitate development of efficient NLP systems making their widespread adoption in real-world applications possible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du 等ICLR 2024 · 被引用 153 次
- Language Model Cascades: Token-Level Uncertainty And BeyondNeha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat 等ICLR 2024 · 被引用 119 次
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja 等ICLR 2026 · 被引用 99 次
- When Does Confidence-Based Cascade Deferral Suffice?Wittawat Jitkrittum, Neha Gupta, Aditya Krishna Menon, Harikrishna Narasimhan 等NeurIPS 2023 · 被引用 76 次
- Online Cascade Learning for Efficient Inference over StreamsLunyiu Nie, Zhimin Ding, Erdong Hu, Christopher M. Jermaine 等ICML 2024 · 被引用 20 次
它引用的顶会 Paper17
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma 等AAAI 2020 · 被引用 656 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
相关 Paper
- Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware DeferralAntónio Farinhas, Nuno Miguel Guerreiro, Sweta Agrawal, Ricardo Rei 等EMNLP 2025 · 被引用 2 次
- Task Cascades for Efficient Unstructured Data ProcessingShreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranSIGMOD 2026 · 被引用 7 次
- The Cascade Transformer: an Application for Efficient Answer Sentence SelectionLuca Soldaini, Alessandro MoschittiACL 2020 · 被引用 3 次
- Window-Based Early-Exit Cascades for Uncertainty Estimation: When Deep Ensembles are More Efficient than Single ModelsGuoxuan Xia, Christos-Savvas BouganisICCV 2023 · 被引用 17 次
- Cut Costs, Not Accuracy: LLM-Powered Data Processing with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranSIGMOD 2026 · 被引用 1 次
