MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
Wenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang, Xindian Ma
Abstract
Quantization significantly accelerates inference in large language models (LLMs) by replacing original high-precision matrices with low-precision counterparts. Recent advances in weight-activation quantization have primarily focused on mapping both weights and activations to the INT4 format. Although the new FP4 Tensor Cores in NVIDIA’s Blackwell architecture offer up to 4 speedup over FP16, existing INT4-based kernels fail to fully exploit this capability due to mismatched data formats. To bridge this gap, we propose MicroMix, a co-designed mixed-precision quantization algorithm and GEMM kernel based on Microscaling (MX) data formats. Tailored for the Blackwell architecture, the MicroMix kernel supports arbitrary combinations of MXFP4, MXFP6, and MXFP8 channels, and produces BFloat16 outputs. To achieve a favorable trade-off between accuracy and efficiency for each linear layer, we introduce quantization thresholds that identify activation elements where lower-precision formats (MXFP4 or MXFP6) incur excessive quantization error. Our algorithm selectively allocates higher-precision channels to preserve accuracy while maintaining compute efficiency. On the Llama and Qwen model families, MicroMix achieves near-FP16 performance across diverse downstream tasks with an average precision of 5 bits. In particular, Qwen2.5-32B-Base, Coder and Math exhibit lossless accuracy on zero-shot, code generation, and mathematical reasoning benchmarks. In addition, on RTX 5070Ti laptop and RTX 5090 GPUs, our kernel achieves 2.29-3.38 acceleration compared to TensorRT-FP16. Our code is available at https://github.com/lwy2020/MicroMix.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d0d5b34-47e6-4c27-a7dc-3eb6dfea1eeaCited by top-tier papers2
- ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMsHaoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu et al.ACL 2026 · 5 citations
- OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference AccelerationXueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu et al.ISCA 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu et al.ASPLOS 2025 · 5 citations
- MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block RepresentationsJiaxiang Zou, Yonghao Chen, Ruilong WU, Xinyu ChenICML 2026
- Quartet: Native FP4 Training Can Be Optimal for Large Language ModelsRoberto L. Castro, Andrei Panferov, Rush Tabesh, Oliver Sieberling et al.NeurIPS 2025 · 38 citations
- INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization FormatsMengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan et al.ICML 2026 · 21 citations
- OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language ModelsJahyun Koo, Dahoon Park, Sangwoo Jung, Jaeha KungDAC 2024 · 12 citations
