ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
Zirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang, Yue Cheng
摘要
Modern model hubs, such as Hugging Face, store tens of petabytes of LLMs, with fine-tuned variants vastly outnumbering base models and dominating storage consumption. Existing storage reduction techniques-such as deduplication and compression-are either LLM-oblivious or not compatible with each other, limiting data reduction effectiveness.
Our large-scale characterization study across all publicly available Hugging Face LLM repositories reveals several key insights: (1) fine-tuned models within the same family exhibit highly structured, sparse parameter differences suitable for delta compression; (2) bitwise similarity enables LLM family clustering; and (3) tensor-level deduplication is better aligned with model storage workloads, achieving high data reduction with low metadata overhead. Building on these insights, we design BitX, an effective, fast, lossless delta compression algorithm that compresses the XORed difference between fine-tuned and base LLMs. We build ZipLLM, a model storage reduction pipeline that unifies tensor-level deduplication and lossless BitX compression. By synergizing deduplication and compression around LLM family clustering, ZipLLM reduces model storage consumption by 54%, over 20% higher than state-of-the-art deduplication and compression approaches. Table 1: Comparison of model storage reduction techniques. Note that existing solutions are limited to use either deduplication or compression. Solution Compression Deduplication Cross-model Throughput Storage Reduction Cons & Pros HuggingFace Xet [79] No Yes Yes Low High No compression support ELF [70] Yes No No High High Lossy compression ZipNN [30] Yes No No Medium Medium Ignores cross-model redundancy FM-Delta [58] Yes No Yes Low Medium Requires identical model structure; lacks BF16 support ZipLLM (ours) Yes Yes Yes High High Lossless and model structure-aware dedup and compression
family structure impacts storage redundancy and compression effectiveness.
• We introduce a novel metric, bit distance, to quantify the similarity between fine-tuned models and their base models.
• Building on this, we design BitX, a highly effective, fast, lossless delta compression algorithm that compresses LLM variants by encoding XOR-based deltas.
• We identify a new ML system design principle: for modern model storage systems, deduplication and lossless compression must be co-designed and unified to fully exploit model structure and redundancy.
• We build ZipLLM, a model storage reduction pipeline that synergizes tensor-level deduplication and lossless BitX compression, achieving higher storage savings for large-scale LLM repositories. Evaluation results show that ZipLLM reduces the storage size of 3,048 sampled LLMs by 54.1%, 20% higher than the state-of-the-art methods. Meanwhile, ZipLLM achieves 2× higher compression throughput (Figure 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight CompressionTim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev 等ICLR 2024 · 被引用 392 次
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 被引用 65 次
相关 Paper
- TensorDex: A Compact, Tensor-Centric Storage System for Modern AI ModelsTingfeng Lan, Zirui Wang, Yunjia Zheng, Zhaoyuan Su 等SOSP 2026
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu 等NeurIPS 2024 · 被引用 10 次
- Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to AskZhaoyuan Su, Ammar Ahmed, Zirui Wang, Ali Anwar 等VLDB 2024
- ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless CompressionRuibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li 等ASPLOS 2026
- Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionXiaohui Wang, Peng Ye, Chenyu Huang, Shenghe Zheng 等NeurIPS 2025
