QStore: Quantization-Aware Compressed Model Storage
Raunak Shah, Zhaoheng Li, Yongjoo Park
摘要
Modern applications commonly leverage large, multi-modal foundation models, in complex workflows that demand the storage and usage of similar models in multiple precisions. A straightforward approach is to maintain a separate file for each model precision (e.g., INT8, BF16), which is indeed taken by model providers such as HuggingFace and Ollama. However, this approach incurs excessive storage costs as a higher precision model (e.g., BF16) is a superset of a lower precision model (e.g., INT8) in terms of information. Unfortunately, simply maintaining only the higher-precision model and requiring every user to dynamically convert the model precision is not desirable because every user of lower precision models must pay the cost for model download and precision conversion.
In this paper, we present QStore, a unified, lossless compression format for simultaneously storing a model in two (high and low) precisions efficiently. Instead of storing low and high-precision models separately, QStore stores low-precision model and only residual information needed to reconstruct high-precision models. The residual information size is significantly smaller than the original high-precision models, thus, achieving high storage cost savings. Moreover, QStore does not compromise model loading speed: The low-precision models can still be loaded quickly, while the high-precision models can also be reconstructed efficiently by merging low-precision data and the residual with QStore's lightweight decoding. We evaluate QStore for compressing multiple precisions of popular foundation models, and show that QStore reduces overall storage cost by up to 2.2× while enabling up to 1.7× and 1.8× faster model saving and loading versus existing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Chipmink: Efficient Delta Identification for Massive Object GraphsSupawit Chockchowwat, Sumay Thakurdesai, Zhaoheng Li, Matthew Krafczyk 等VLDB 2026 · 被引用 1 次
- stratum: A System Infrastructure for Massive Agent-Centric ML WorkloadsArnab Phani, Elias Strauss, Sebastian SchelterVLDB 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
相关 Paper
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu 等NeurIPS 2024 · 被引用 10 次
- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMsYeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim 等ICML 2024 · 被引用 51 次
- NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMsHaeun Lee, Omin Kwon, Yeonhong Park, Jae W. LeeNeurIPS 2025 · 被引用 5 次
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang 等NSDI 2026 · 被引用 8 次
- Matryoshka QuantizationPranav Ajit Nair, Puranjay Datta, Jeff Dean, Prateek Jain 等ICML 2025
