Lune

NSDI2026Top-tier venue

ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression

Zirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang, Yue Cheng

2026Year
8Citations

Abstract

Modern model hubs, such as Hugging Face, store tens of petabytes of LLMs, with fine-tuned variants vastly outnumbering base models and dominating storage consumption. Existing storage reduction techniques-such as deduplication and compression-are either LLM-oblivious or not compatible with each other, limiting data reduction effectiveness.

Our large-scale characterization study across all publicly available Hugging Face LLM repositories reveals several key insights: (1) fine-tuned models within the same family exhibit highly structured, sparse parameter differences suitable for delta compression; (2) bitwise similarity enables LLM family clustering; and (3) tensor-level deduplication is better aligned with model storage workloads, achieving high data reduction with low metadata overhead. Building on these insights, we design BitX, an effective, fast, lossless delta compression algorithm that compresses the XORed difference between fine-tuned and base LLMs. We build ZipLLM, a model storage reduction pipeline that unifies tensor-level deduplication and lossless BitX compression. By synergizing deduplication and compression around LLM family clustering, ZipLLM reduces model storage consumption by 54%, over 20% higher than state-of-the-art deduplication and compression approaches. Table 1: Comparison of model storage reduction techniques. Note that existing solutions are limited to use either deduplication or compression. Solution Compression Deduplication Cross-model Throughput Storage Reduction Cons & Pros HuggingFace Xet [79] No Yes Yes Low High No compression support ELF [70] Yes No No High High Lossy compression ZipNN [30] Yes No No Medium Medium Ignores cross-model redundancy FM-Delta [58] Yes No Yes Low Medium Requires identical model structure; lacks BF16 support ZipLLM (ours) Yes Yes Yes High High Lossless and model structure-aware dedup and compression

family structure impacts storage redundancy and compression effectiveness.

• We introduce a novel metric, bit distance, to quantify the similarity between fine-tuned models and their base models.

• Building on this, we design BitX, a highly effective, fast, lossless delta compression algorithm that compresses LLM variants by encoding XOR-based deltas.

• We identify a new ML system design principle: for modern model storage systems, deduplication and lossless compression must be co-designed and unified to fully exploit model structure and redundancy.

• We build ZipLLM, a model storage reduction pipeline that synergizes tensor-level deduplication and lossless BitX compression, achieving higher storage savings for large-scale LLM repositories. Evaluation results show that ZipLLM reduces the storage size of 3,048 sampled LLMs by 54.1%, 20% higher than the state-of-the-art methods. Meanwhile, ZipLLM achieves 2× higher compression throughput (Figure 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b065ab8e-b040-4964-be3d-0af381da7ee1

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines