OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, Eunhyeok Park
Abstract
Large language models (LLMs) with hundreds of billions of parameters require powerful server-grade GPUs for inference, limiting their practical deployment. To address this challenge, we introduce the outlier-aware weight quantization (OWQ) method, which aims to minimize LLM's footprint through low-precision representation. OWQ prioritizes a small subset of structured weights sensitive to quantization, storing them in high-precision, while applying highly tuned quantization to the remaining dense weights. This sensitivity-aware mixed-precision scheme reduces the quantization error notably, and extensive experiments demonstrate that 3.1-bit models using OWQ perform comparably to 4-bit models optimized by OPTQ. Furthermore, OWQ incorporates a parameter-efficient fine-tuning for task-specific adaptation, called weak column tuning (WCT), enabling accurate task-specific LLM adaptation with minimal memory overhead in the optimized format. OWQ represents a notable advancement in the flexibility, efficiency, and practicality of LLM optimization literature. The source code is available at https://github.com/xvyaward/owq.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers42
- PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM CompressionVladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev et al.NeurIPS 2024 · 69 citations
- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMsYeonhong Park, Jake Hyun, SangLyul Cho, Bonggeun Sim et al.ICML 2024 · 51 citations
- APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language ModelsZiyi Guan, Hantao Huang, Yupeng Su, Hong Huang et al.DAC 2024 · 26 citations
- DartQuant: Efficient Rotational Distribution Calibration for LLM QuantizationYuantian Shao, Yuanteng Chen, Peisong Wang, Jianlin Yu et al.NeurIPS 2025 · 20 citations
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM InferenceYesheng Liang, Haisheng Chen, Song Han, Zhijian LiuICLR 2026 · 19 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
Related papers
- Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and VariationsPatrick Blumenberg, Thomas Graave, Tim FingscheidtICLR 2026 · 5 citations
- Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer QuantizationJeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park et al.NeurIPS 2023 · 157 citations
- RoMeo: Mitigating Dual-dimensional Outliers with Rotated Mixed Precision QuantizationQihao Zhang, Mingliang Tang, Mingshu Zhai, Kinman Lei et al.PPoPP 2026
- Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient AdaptationJingjing Xie, Yuxin Zhang, Mingbao Lin, Liujuan Cao et al.ACM MM 2024 · 5 citations
- Zeroth-Order Fine-Tuning of LLMs with Transferable Static SparsityWentao Guo, Jikai Long, Yimeng Zeng, Zirui Liu et al.ICLR 2025
