Task Vector Quantization for Memory-Efficient Model Merging
Youngeun Kim, Seunghwan Lee, Aecheon Jung, Bogon Ryu, Sungeun Hong
摘要
Model merging enables efficient multi-task models by combining task-specific fine-tuned checkpoints. However, storing multiple task-specific checkpoints requires significant memory, limiting scalability and restricting model merging to larger models and diverse tasks. In this paper, we propose quantizing task vectors (i.e., the difference between pre-trained and fine-tuned checkpoints) instead of quantizing fine-tuned checkpoints. We observe that task vectors exhibit a narrow weight range, enabling lowprecision quantization ( bit) within existing task vector merging frameworks. To further mitigate quantization errors within ultra-low bit precision (e.g., 2 bit), we introduce Residual Task Vector Quantization, which decomposes the task vector into a base vector and offset component. We allocate bits based on quantization sensitivity, ensuring precision while minimizing error within a memory budget. Experiments on image classification and dense prediction show our method maintains or improves model merging performance while using only 8 % of the memory required for full-precision checkpoints. Our code is available at https://aim-skku.github.io/TVQ/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SyMerge: From Non-Interference to Synergistic Merging via Single-Layer AdaptationAecheon Jung, Seunghwan Lee, Dongyoon Han, Sungeun HongICML 2026 · 被引用 1 次
- ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language ModelsYoungeun Kim, Youjia Zhang, Huiling Liu, Aecheon Jung 等CVPR 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- Less is More: Efficient Model Merging with Binary Task SwitchBiqing Qi, Fangyuan Li, Zhen Wang, Junqi Gao 等CVPR 2025
- CABS: Conflict-Aware and Balanced Sparsification for Enhancing Model MergingZongzhen Yang, Binhang Qi, Hailong Sun, Wenrui Long 等ICML 2025
- BitDelta: Your Fine-Tune May Only Be Worth One BitJames Liu, Guangxuan Xiao, Kai Li, Jason D. Lee 等NeurIPS 2024 · 被引用 50 次
- Task Singular Vectors: Reducing Task Interference in Model MergingAntonio Andrea Gargiulo, Donato Crisostomi, Maria Sofia Bucarelli, Simone Scardapane 等CVPR 2025
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationChun Hu, Junhui He, Shangyu Wu, Yuxin He 等EMNLP 2025 · 被引用 1 次
