MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
Akshat Ramachandran, Souvik Kundu, Tushar Krishna
Abstract
Quantization of foundational models (FMs) is significantly more challenging than traditional DNNs due to the emergence of large magnitude values called outliers. Existing outlier-aware algorithmarchitecture co-design techniques either use mixed-precision, retaining outliers at high precision but compromise hardware efficiency, or quantize inliers and outliers at the same precision, improving hardware efficiency at the cost of accuracy. To address this mutual exclusivity, we propose MicroScopiQ, a novel co-design technique that leverages pruning to complement outlier-aware quantization. MicroScopiQ retains outliers at higher precision while pruning a certain fraction of least important weights to distribute the additional outlier bits; ensuring high accuracy, aligned memory and hardware efficiency. We design a high-throughput, low overhead accelerator architecture composed of multi-precision INT processing elements and a network-on-chip called ReCoN that efficiently abstracts the complexity of supporting high-precision outliers. Additionally, unlike prior techniques, MicroScopiQ does not assume any locality of outlier weights, enabling applicability to a broad range of FMs. Extensive experiments across diverse quantization settings demonstrate that MicroScopiQ achieves state-of-the-art quantization accuracy, while delivering up to 3× faster inference and 2× lower energy consumption compared to existing alternatives. Code is available at: MicroScopiQ-LLM-Quantization.git CCS Concepts • Computer systems organization → Systolic arrays; Neural networks; Data flow architectures; • Networks → NoC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ae2dd21-6639-4b5b-b4f9-e34057692104Cited by top-tier papers8
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning ModelsAkshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan et al.ICLR 2026 · 19 citations
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang et al.ICLR 2026 · 12 citations
- Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error ReductionJatin Chhugani, Geonhwa Jeong, Bor-Yiing Su, Yunjie Pan et al.ICML 2026 · 6 citations
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu et al.ISCA 2026 · 4 citations
- GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM InferenceSangjin Kim, Yuseon Chou, Byeongcheol Kim, Jungjun Oh et al.HPCA 2026 · 2 citations
Builds on37
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT ComputationZihan Zou, Shikuang Chen, Chen Zhang, Xing Wang et al.DAC 2025
- DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs AccelerationZhuoquan Yu, Huidong Ji, Yue Cao, Junfu Wu et al.DAC 2025 · 1 citation
- An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language ModelsYingbo Hao, Huangxu Chen, Yi Zou, Yanfeng YangDAC 2025 · 1 citation
- Oltron: Algorithm-Hardware Co-design for Outlier-Aware Quantization of LLMs with Inter-/Intra-Layer AdaptationChenhao Xue, Chen Zhang, Xun Jiang, Zhutianya Gao et al.DAC 2024 · 11 citations
- BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionFangxin Liu, Ning Yang, Zongwu Wang, Xuanpeng Zhu et al.DAC 2025 · 1 citation
