BitDelta: Your Fine-Tune May Only Be Worth One Bit
James Liu, Guangxuan Xiao, Kai Li, Jason D. Lee, Song Han, Tri Dao, Tianle Cai
摘要
Large Language Models (LLMs) are typically trained in two phases: pre-training on large internet-scale datasets, and fine-tuning for downstream tasks. Given the higher computational demand of pre-training, it's intuitive to assume that fine-tuning adds less new information to the model, and is thus more compressible. We explore this assumption by decomposing the weights of fine-tuned models into their pre-trained components and an additional delta. We introduce a simple method, BitDelta, which successfully quantizes this delta down to 1 bit without compromising performance. This interesting finding not only highlights the potential redundancy of information added during fine-tuning, but also has significant implications for the multi-tenant serving and multi-tenant storage of fine-tuned models. By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, which can also be translated to enhanced generation latency in multi-tenant settings. We validate BitDelta through experiments across Llama-2 and Mistral model families, and on models up to 70B parameters, showcasing minimal performance degradation over all tested settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Free-Merging: Fourier Transform for Efficient Model MergingShenghe Zheng, Hongzhi WangICCV 2025 · 被引用 12 次
- Knowledge Fusion of Large Language Models via Modular SkillPacksGuodong Du, Zhuo Li, Xuanning Zhou, Junlin Li 等ICLR 2026 · 被引用 9 次
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang 等NSDI 2026 · 被引用 8 次
- DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMsXiaozhe Yao, Qinghao Hu, Ana KlimovicEuroSys 2025 · 被引用 6 次
- Personality Vector: Modulating Personality of Large Language Models by Model MergingSeungjong Sun, Seo Yeon Baek, Jang Hyun KimEMNLP 2025 · 被引用 5 次
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language ModelsBowen Ping, Shuo Wang, Hanqing Wang, Xu Han 等NeurIPS 2024 · 被引用 25 次
- Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta CompressionXiaohui Wang, Peng Ye, Chenyu Huang, Shenghe Zheng 等NeurIPS 2025
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 被引用 14 次
- NanoQuant: Efficient Sub-1-Bit Quantization of Large Language ModelsHyochan Chong, Dongkyu Kim, Changdong Kim, Minseop ChoiICML 2026
- FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsWanyi Ning, Jingyu Wang, Qi Qi, Mengde Zhu 等NeurIPS 2024 · 被引用 10 次
