GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
Jinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens J. S. Schaefer, Deokjae Lee, Yeonhong Park, Jae W. Lee, Hyun Oh Song
Abstract
Post-training quantization is a key technique for reducing the memory and inference latency of large language models by quantizing weights and activations without requiring retraining. However, existing methods either (1) fail to account for the varying importance of hidden features to the end loss or, when incorporating end loss, (2) neglect the critical interactions between model weights. To address these limitations, we propose GuidedQuant, a novel quantization approach that integrates gradient information from the end loss into the quantization objective while preserving cross-weight dependencies within output channels. GuidedQuant consistently boosts the performance of state-of-the-art quantization methods across weight-only scalar, weight-only vector, and weight-and-activation quantization. Additionally, we introduce a novel non-uniform scalar quantization algorithm, which is guaranteed to monotonically decrease the quantization objective value, and outperforms existing methods in this category. We release the code at https://github. com/snu-mllab/GuidedQuant .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc216d8d-c55a-4ef0-ac7f-f5c1804fcf42Cited by top-tier papers3
- Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM DeploymentDeokjae Lee, Hyun Oh SongNeurIPS 2025 · 9 citations
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMsGunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim et al.ICLR 2026 · 8 citations
- GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMsSelim An, Il hong Suh, Yeseong KimICLR 2026
Builds on17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- LQER: Low-Rank Quantization Error Reconstruction for LLMsCheng Zhang, Jianyi Cheng, George Anthony Constantinides, Yiren ZhaoICML 2024 · 33 citations
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li et al.ICML 2025
- SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language ModelsHan Liu, Haotian Gao, Xiaotong Zhang, Changya Li et al.KDD 2025 · 1 citation
- EasyQuant: An Efficient Data-free Quantization Algorithm for LLMsHanlin Tang, Yifu Sun, Decheng Wu, Kai Liu et al.EMNLP 2023 · 4 citations
- TurboBoA: Faster and Exact Attention-aware Quantization without BackpropagationJunhan Kim, Yeo Jeong Park, Seungwoo Son, Chungman Lee et al.ICLR 2026 · 3 citations
