Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference
Suyi Li, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng, Xusheng Chen, Yizhou Shan, Binhang Yuan, Wei Wang
Abstract
Low-Rank Adaptation (LoRA) is a popular approach that adapts a base large language model (LLM) to domain-specific tasks by adding lightweight trainable adapters. In this paper, we present TOPPINGS, a system that efficiently serves many LoRA adapters derived from a common base LLM. TOP-PINGS pins the base model on GPUs and dynamically loads the requested LoRA adapters from host memory as new requests arrive. In view of the high GPU loading overhead, which not only delays the time-to-first-token of the newly arrived request but also interrupts the ongoing decoding of all inflight queries when continuous batching is in use, TOP-PINGS proposes a CPU-assisted LoRA serving approach. It simultaneously uses CPUs to compute the lightweight adaption for prefilling as the requested LoRA adapter is being loaded onto GPUs; it then switches to the GPUs after loading completes to resume the remaining computation. TOPPINGS develops a highly optimized synchronization mechanism and pipeline loading scheme to efficiently coordinate LoRA computation on the CPUs and GPUs. TOPPINGS further designs a rank-aware scheduling algorithm that optimally schedules heterogeneous LoRA requests to maximize the SLO attainment. Compared with the state-of-the-art LoRA serving systems, TOPPINGS improves the average request serving latency by up to 1.7× and achieves an SLO attainment of up to 99%.
128 256 512 No. of adapters in a server 40 60 Time per Token (ms) GPU OOM Cached S-LoRA dLoRA Toppings
- "Topping" is a layer of food added to another food base to enhance flavor. Here, it metaphorically refers to a LoRA that adapts a base LLM's outputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85245b2c-1d71-4e83-b4f2-0f54548539cbCited by top-tier papers3
- Katz: Efficient Workflow Serving for Diffusion Models with Many AdaptersSuyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu et al.USENIX ATC 2025 · 14 citations
- Kairox: Adaptive GPU-CPU Hybrid LLM Inference via Online Neuron BalancingYapeng Jiang, Minghao Gan, Zicong Hong, Wuhui Chen et al.OSDI 2026
- ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingJiuchen Shi, Hang Zhang, Yixiao Wang, Quan Chen et al.HPCA 2026
Builds on26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- DyMerge-LoRA: On-GPU Post-Merge Fusion for High-Throughput Multi-Tenant Composite LoRA ServingRui Xu, Long Chen, Huazheng Lao, Jinquan Zhang et al.KDD 2026
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- PLoRA: Efficient Concurrent LoRA Training for Large Language ModelsMinghao Yan, Zhuang Wang, Zhen Jia, Shivaram Venkataraman et al.ICML 2026 · 5 citations
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 1 citation
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference EnvironmentsNikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu et al.MICRO 2025 · 3 citations
