Tabi: An Efficient Multi-Level Inference System for Large Language Models
Yiding Wang, Kai Chen, Haisheng Tan, Kun Guo
Abstract
Today's trend of building ever larger language models (LLMs), while pushing the performance of natural language processing, adds significant latency to the inference stage. We observe that due to the diminishing returns of adding parameters to LLMs, a smaller model could make the same prediction as a costly LLM for a majority of queries. Based on this observation, we design Tabi, an inference system with a multi-level inference engine that serves queries using small models and optional LLMs for demanding applications. Tabi is optimized for discriminative models (i.e., not generative LLMs) in a serving framework. Tabi uses the calibrated confidence score to decide whether to return the accurate results of small models extremely fast or re-route them to LLMs. For re-routed queries, it uses attention-based word pruning and weighted ensemble techniques to offset the system overhead and accuracy loss. We implement and evaluate Tabi with multiple tasks and models. Our result shows that Tabi achieves 21%-40% average latency reduction (with comparable tail latency) over the state-of-the-art while meeting LLM-grade high accuracy targets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang et al.USENIX ATC 2024 · 273 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPUYixin Song, Zeyu Mi, Haotong Xie, Haibo ChenSOSP 2024 · 86 citations
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- Approximate Caching for Efficiently Serving Text-to-Image Diffusion ModelsShubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam et al.NSDI 2024 · 44 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
Related papers
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation SystemsLiang Wang, Kai Wang, Ranjun Jia, Kai Lu et al.ICDE 2026
- Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMsRunyu Peng, Yunhua Zhou, Kai Lv, Yang Gao et al.EMNLP 2025
- SATER: A Self-Aware and Token-Efficient Approach to Routing and CascadingYuanzhe Shen, Yide Liu, Zisu Huang, Ruicheng Yin et al.EMNLP 2025
- Towards Optimal Caching and Model Selection for Large Model InferenceBanghua Zhu, Ying Sheng, Lianmin Zheng, Clark W. Barrett et al.NeurIPS 2023 · 19 citations
