ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
Zirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang, Yue Cheng
Abstract
Modern model hubs, such as Hugging Face, store tens of petabytes of LLMs, with fine-tuned variants vastly outnumbering base models and dominating storage consumption. Existing storage reduction techniques-such as deduplication and compression-are either LLM-oblivious or not compatible with each other, limiting data reduction effectiveness.
Our large-scale characterization study across all publicly available Hugging Face LLM repositories reveals several key insights: (1) fine-tuned models within the same family exhibit highly structured, sparse parameter differences suitable for delta compression; (2) bitwise similarity enables LLM family clustering; and (3) tensor-level deduplication is better aligned with model storage workloads, achieving high data reduction with low metadata overhead. Building on these insights, we design BitX, an effective, fast, lossless delta compression algorithm that compresses the XORed difference between fine-tuned and base LLMs. We build ZipLLM, a model storage reduction pipeline that unifies tensor-level deduplication and lossless BitX compression. By synergizing deduplication and compression around LLM family clustering, ZipLLM reduces model storage consumption by 54%, over 20% higher than state-of-the-art deduplication and compression approaches. Table 1: Comparison of model storage reduction techniques. Note that existing solutions are limited to use either deduplication or compression. Solution Compression Deduplication Cross-model Throughput Storage Reduction Cons & Pros HuggingFace Xet [79] No Yes Yes Low High No compression support ELF [70] Yes No No High High Lossy compression ZipNN [30] Yes No No Medium Medium Ignores cross-model redundancy FM-Delta [58] Yes No Yes Low Medium Requires identical model structure; lacks BF16 support ZipLLM (ours) Yes Yes Yes High High Lossless and model structure-aware dedup and compression
family structure impacts storage redundancy and compression effectiveness.
• We introduce a novel metric, bit distance, to quantify the similarity between fine-tuned models and their base models.
• Building on this, we design BitX, a highly effective, fast, lossless delta compression algorithm that compresses LLM variants by encoding XOR-based deltas.
• We identify a new ML system design principle: for modern model storage systems, deduplication and lossless compression must be co-designed and unified to fully exploit model structure and redundancy.
• We build ZipLLM, a model storage reduction pipeline that synergizes tensor-level deduplication and lossless BitX compression, achieving higher storage savings for large-scale LLM repositories. Evaluation results show that ZipLLM reduces the storage size of 3,048 sampled LLMs by 54.1%, 20% higher than the state-of-the-art methods. Meanwhile, ZipLLM achieves 2× higher compression throughput (Figure 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b065ab8e-b040-4964-be3d-0af381da7ee1Builds on15
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight CompressionTim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev et al.ICLR 2024 · 392 citations
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 65 citations
Related papers
- TensorDex: A Compact, Tensor-Centric Storage System for Modern AI ModelsTingfeng Lan, Zirui Wang, Yunjia Zheng, Zhaoyuan Su et al.SOSP 2026
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu et al.NeurIPS 2024 · 10 citations
- Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to AskZhaoyuan Su, Ammar Ahmed, Zirui Wang, Ali Anwar et al.VLDB 2024
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li et al.ASPLOS 2026
- Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionXiaohui Wang, Peng Ye, Chenyu Huang, Shenghe Zheng et al.NeurIPS 2025
