Towards Quantization-Aware Training for Ultra-Low-Bit Reasoning LLMs
Yasuyuki Okoshi, Hikari Otsuka, Daichi Fujiki, Masato Motomura
Abstract
Large language models (LLMs) have achieved remarkable performance across diverse reasoning tasks, yet their deployment is hindered by prohibitive computational and memory costs. Quantization-aware training (QAT) enables ultralow-bit compression (< 4 bits per weight), but existing QAT methods often degrade reasoning capability, partly because complex knowledge structures are introduced during the post-training process in LLMs. In this paper, through a systematic investigation of how quantization affects different data domains, we find that its impact on pre-training and reasoning capabilities differs. Building on this insight, we propose a novel two-stage QAT pipeline specifically designed for reasoning LLMs. In the first stage, we quantize the model using mixed-domain calibration data to preserve essential capabilities across domains; in the second stage, we fine-tune the quantized model with a teacher-guided reward-rectification loss to restore reasoning capability. We first demonstrate that mixed-domain calibration outperforms single-domain calibration at maximum 2.74% improvement on average over six tasks including reasoning and pretrained tasks. Following experiments on five reasoning benchmarks show that our 2-bit-quantized Qwen3-8B outperforms post-training quantization (PTQ) baselines by 50.45% on average. Moreover, compared to ultra-low-bit-specialized models such as BitNet-2B4T, our pipeline achieves approximately 2% higher mathematical-reasoning accuracy with fewer than 1B tokens. Code is available: https://github.com/yasu0001/ReasoningQAT . * Equal contribution REASONING-ORIENTED TWO-STAGE QUANTIZATION AWARE TRAINING This section proposes a novel QAT pipeline that enables the preservation of reasoning capabilities after ultra-low-bit quantization. We first analyze the impact of quantization on various knowledge domains, and based on these findings, we introduce the reasoning-oriented QAT pipeline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79b87ca9-3409-4ac1-ab79-69a55f25304aBuilds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
Related papers
- PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMsTengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao et al.ICLR 2026 · 18 citations
- PB-LLM: Partially Binarized Large Language ModelsZhihang Yuan, Yuzhang Shang, Zhen DongICLR 2024 · 91 citations
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training DataSyeda Nahida Akter, Shrimai Prabhumoye, Eric Nyberg, Mostofa Patwary et al.ICLR 2026 · 27 citations
- Scaling Law for Quantization-Aware TrainingMengzhao Chen, Chaoyi Zhang, Jing Liu, Zeng et al.ICML 2026 · 16 citations
- EfficientQAT: Efficient Quantization-Aware Training for Large Language ModelsMengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang et al.ACL 2025
