DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
Xiaozhe Yao, Qinghao Hu, Ana Klimovic
Abstract
Fine-tuning large language models (LLMs) greatly improves model quality for downstream tasks. However, serving many fine-tuned LLMs concurrently is challenging due to the sporadic, bursty, and varying request patterns of different LLMs. To bridge this gap, we present DeltaZip, an LLM serving system that efficiently serves multiple full-parameter fine-tuned models concurrently by aggressively compressing model deltas by up to 10× while maintaining high model quality.
The key insight behind this design is that fine-tuning results in small-magnitude changes to the pre-trained model. By co-designing the serving system with the compression algorithm, DeltaZip achieves 2× to 12× improvement in throughput compared to the state-of-the-art systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 961c56ea-5629-4e51-89e6-33dd99c5fca8Cited by top-tier papers6
- Resource Multiplexing in Tuning and Serving Large Language ModelsYongjun He, Haofeng Yang, Yao Lu, Ana Klimovic et al.USENIX ATC 2025 · 9 citations
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang et al.NSDI 2026 · 8 citations
- RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMsYongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu et al.NSDI 2026 · 4 citations
- ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMsYan Yang, Yixia Li, Hongru Wang, Xuetao Wei et al.ACL 2025 · 4 citations
- Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionXiaohui Wang, Peng Ye, Chenyu Huang, Shenghe Zheng et al.NeurIPS 2025
Builds on34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Compress then Serve: Serving Thousands of LoRA Adapters with Little OverheadRickard Brüel Gabrielsson, Jiacheng Zhu, Onkar Bhardwaj, Leshem Choshen et al.ICML 2025
- Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language ModelsBowen Ping, Shuo Wang, Hanqing Wang, Xu Han et al.NeurIPS 2024 · 25 citations
- Efficient Multi-task LLM Quantization and Serving for Multiple LoRA AdaptersYifei Xia, Fangcheng Fu, Wentao Zhang, Jiawei Jiang et al.NeurIPS 2024 · 25 citations
- MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM ServingJiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li et al.ICML 2024 · 51 citations
- High Throughput and Low Latency LLM Serving via Adaptive KV CachingWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2026
