ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training
Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi
Abstract
Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints. Microscaled FP4 formats enable efficient FP4 deployment; however, fully quantizing weights, activations, and KV caches (W4A4KV4) causes severe reasoning degradation that existing PTQ and QAT fail to recover. We identify that FP4 failures concentrate on low-entropy tokens—precise symbolic commitments such as digits and operators—where quantization noise inflates sampling errors that cascade through reasoning traces. Based on this insight, we propose ReQAT, a reasoning-centric FP4 training framework with three components: (i) Trace-Aligned QAT (TAQ), which revisits identical reasoning traces to focus updates on critical low-entropy decisions; (ii) Selective Entropy Minimization (SEM), which reinforces confidence at low-entropy positions; and (iii) Q-FIT, a quantization-friendly initialization that jointly calibrates RoPE-consistent KV cache transformations to stabilize QAT. Under the same training budget, ReQAT not only recovers but surpasses BF16 fine-tuning accuracy—achieving while delivering up to throughput speedup on NVIDIA DGX Spark and on B200. This is the first demonstration that FP4 QAT can exceed full-precision accuracy for LRMs with over 3× speedup on production hardware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1716d61d-aa80-4bee-8173-3a2e833a13ebBuilds on15
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong et al.ICML 2024 · 436 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han et al.NeurIPS 2025 · 185 citations
Related papers
- Attn-QAT: 4-Bit Attention With Quantization-Aware TrainingPeiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang et al.ICML 2026 · 2 citations
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM ReasoningJiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon KimNeurIPS 2025 · 25 citations
- Training Large Reasoning Models Efficiently via Progressive Thought EncodingZeliang Zhang, Xiaodong Liu, Hao Cheng, Hao Sun et al.ICLR 2026 · 2 citations
- Bridging the Gap Between Promise and Performance for Microscaling FP4 QuantizationVage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov et al.ICLR 2026 · 38 citations
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning ModelsJunhyuck Kim, Ethan Ewer, Taehong Moon, Jongho Park et al.ICLR 2026 · 2 citations
