1-Bit Wonder: Improving QAT Performance in the Low-Bit Regime through K-Means Quantization
Sohir Maskey, Constantin Eichenberg, Johannes Messner, Douglas Orr
摘要
Quantization-aware training (QAT) is an effective method to drastically reduce the memory footprint of LLMs while keeping performance degradation at an acceptable level. However, the optimal choice of quantization format and bit-width presents a challenge in practice. The full design space of quantization is not fully explored in the context of QAT, and the precise trade-off between quantization and downstream performance is poorly understood, as comparisons often rely solely on perplexity-based evaluations. In this work, we address these shortcomings with an empirical study of QAT in the low-bit regime. We show that k-means based weight quantization outperforms integer formats and can be implemented efficiently on standard hardware. Furthermore, we find that, under a fixed inference memory budget, the best performance on generative downstream tasks is achieved with -bit quantized weights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong 等ICML 2024 · 被引用 306 次
- Evaluating Quantized Large Language ModelsShiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu 等ICML 2024 · 被引用 88 次
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining DatasetDan Su, Kezhi Kong, Ying Lin, Joseph Jennings 等ACL 2025
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon 等ICLR 2025
相关 Paper
- LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector QuantizationHaoyu Wang, Xingyu Yu, Haiyan Zhao, Fengxiang Wang 等ICML 2026 · 被引用 1 次
- EfficientQAT: Efficient Quantization-Aware Training for Large Language ModelsMengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang 等ACL 2025
- QuEST: Stable Training of LLMs with 1-Bit Weights and ActivationsAndrei Panferov, Jiale Chen, Soroush Tabesh, Mahdi Nikdan 等ICML 2025
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 被引用 14 次
- OneBit: Towards Extremely Low-bit Large Language ModelsYuzhuang Xu, Xu Han, Zonghan Yang, Shuo Wang 等NeurIPS 2024 · 被引用 110 次
