Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
Hong Huang, Dapeng Wu
摘要
Large language models (LLMs) have made exciting achievements across various domains, yet their deployment on resource-constrained personal devices remains hindered by the prohibitive computational and memory demands of task-specific fine-tuning. While quantization offers a pathway to efficiency, existing methods struggle to balance performance and overhead, either incurring high computational/memory costs or failing to address activation outliers, a critical bottleneck in quantized fine-tuning. To address these challenges, we propose the Outlier Spatial Stability Hypothesis (OSSH): During fine-tuning, certain activation outlier channels retain stable spatial positions across training iterations. Building on OSSH, we propose Quaff, a Quantized parameter-efficient finetuning framework for LLMs, optimizing lowprecision activation representations through targeted momentum scaling. Quaff dynamically suppresses outliers exclusively in invariant channels using lightweight operations, eliminating full-precision weight storage and global rescaling while reducing quantization errors. Extensive experiments across ten benchmarks validate OSSH and demonstrate Quaff's efficacy. Specifically, on the GPQA reasoning benchmark, Quaff achieves a 1.73× latency reduction and 30% memory savings over full-precision fine-tuning while improving accuracy by 0.6% on the Phi-3 model, reconciling the triple trade-off between efficiency, performance, and deployability. By enabling consumer-grade GPU fine-tuning (e.g., RTX 2080 Super) without sacrificing model utility, Quaff democratizes personalized LLM deployment. The code is available at https: //github.com/Little0o0/Quaff.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Tequila: Trapping-free Ternary Quantization for Large Language ModelsHong Huang, Decheng Wu, Rui Cen, Guanghua Yu 等ICLR 2026 · 被引用 12 次
- FedRTS: Federated Robust Pruning via Combinatorial Thompson SamplingHong Huang, Jinhai Yang, Yuan Chen, Jiaxun Ye 等NeurIPS 2025 · 被引用 7 次
- REDOUBT: Duo Safety Validation for Autonomous Vehicle Motion PlanningShuguang Wang, Qian Zhou, Kui Wu, Dapeng Wu 等NeurIPS 2025 · 被引用 6 次
- AE: Towards Compositional Model EditingHongming Piao, Hao Wang, Dapeng Wu, Ying WeiNeurIPS 2025 · 被引用 3 次
- Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained SparsificationHong Huang, Decheng Wu, Qiangqiang Hu, Guanghua Yu 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
相关 Paper
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 被引用 19 次
- DecDEC: A Systems Approach to Advancing Low-Bit LLM QuantizationYeonhong Park, Jake Hyun, Hojoon Kim, Jae W. LeeOSDI 2025 · 被引用 9 次
- OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language ModelsChanghun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim 等AAAI 2024 · 被引用 134 次
- LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware GridTianyi Zhang, Anshumali ShrivastavaICLR 2025
- QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language ModelsJing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong 等ICLR 2024 · 被引用 75 次
