Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask
Zhaoyuan Su, Ammar Ahmed, Zirui Wang, Ali Anwar, Yue Cheng
摘要
As the number of pre-trained machine learning (ML) models is growing exponentially, data reduction tools are not catching up. Existing data reduction techniques are not specifically designed for pre-trained model (PTM) dataset files. This is largely due to a lack of understanding of the patterns and characteristics of these datasets, especially those relevant to data reduction and compressibility. This paper presents the first, exhaustive analysis to date of PTM datasets on storage compressibility. Our analysis spans different types of data reduction and compression techniques, from hashbased data deduplication, data similarity detection, to dictionarycoding compression. Our analysis explores these techniques at three data granularity levels, from model layers, model chunks, to model parameters. We draw new observations that indicate that modern data reduction tools are not effective when handling PTM datasets. There is a pressing need for new compression methods that take into account PTMs' data characteristics for effective storage reduction. Motivated by our findings, we design Elf, a simple yet effective, error-bounded, lossy floating-point compression method. Elf transforms floating-point parameters in such a way that the common exponent field of the transformed parameters can be completely eliminated to save storage space. We develop Elves, a compression framework that integrates Elf along with several other data reduction methods. Elves uses the most effective method to compress PTMs that exhibit different patterns. Evaluation shows that Elves achieves an overall compression ratio of 1.52×, which is 1.31×, 1.32× and 1.29× higher than a general-purpose compressor (zstd), an error-bounded lossy compressor (SZ3), and the uniform model quantization, respectively, with negligible model accuracy loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inferenceYushu Zhao, Zheng Wang, Minjia ZhangICML 2026 · 被引用 8 次
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang 等NSDI 2026 · 被引用 8 次
- NeurStore: Efficient In-database Deep Learning Model Management SystemSiqi Xiang, Sheng Wang, Xiaokui Xiao, Cong Yue 等SIGMOD 2026 · 被引用 2 次
- LiquidCache: Efficient Pushdown Caching for Cloud-Native Data AnalyticsXiangpeng Hao, Andrew Lamb, Yibo Wu, Andrea C. Arpaci-Dusseau 等VLDB 2025
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Optimizing Error-Bounded Lossy Compression for Scientific Data by Dynamic Spline InterpolationKai Zhao, Sheng Di, Maxim Dmitriev, Thierry-Laurent D. Tonellot 等ICDE 2021 · 被引用 151 次
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 被引用 88 次
- Significantly Improving Lossy Compression for HPC Datasets with Second-Order Prediction and Parameter OptimizationKai Zhao, Sheng Di, Xin Liang, Sihuan Li 等HPDC 2020 · 被引用 77 次
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
相关 Paper
- Elf: Erasing-based Lossless Floating-Point CompressionRuiyuan Li, Zheng Li, Yi Wu, Chao Chen 等VLDB 2023 · 被引用 44 次
- Camel: Efficient Compression of Floating-Point Time SeriesYuanyuan Yao, Lu Chen, Ziquan Fang, Yunjun Gao 等SIGMOD 2025 · 被引用 4 次
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu 等NeurIPS 2024 · 被引用 10 次
- AWARE: Workload-aware, Redundancy-exploiting Linear AlgebraSebastian Baunsgaard, Matthias BoehmSIGMOD 2023 · 被引用 4 次
- ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint ShrinkingWenshuo Li, Xinghao Chen, Han Shu, Yehui Tang 等ICML 2024 · 被引用 11 次
