Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation
Huachen Qi, Ruiyu Zhuo, Bowen Shi, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen
Abstract
Large Language Models continue to scale in size and capability, driving substantial computational and memory demands. Mixtureof-Experts (MoE) architectures alleviate this cost by activating only a sparse subset of experts per token, enabling efficient scaling without proportional increases in inference compute. However, quantization in MoE models remains challenging due to heterogeneous sensitivity across experts and their internal linear layers. Existing mixed-precision frameworks such as Mixed-precision Quantization for MoE (MxMoE) require full quantization-loss evaluation for expert-layer-and-bit configurations, incurring prohibitive profiling cost. To address this, we propose FRI-MxMoE, a profilingfree mixed-precision quantization framework that reformulates MoE calibration from exhaustive expert-wise profiling to sparse anchor profiling followed by Fuzzy Rule Interpolation. By constructing a fuzzy rule base in the intraexpert layer feature space (bit-width, activation variance, parameter scale), our method predicts quantization error from only sparse samples while remaining compatible with existing mixed-precision allocation objectives. Extensive experiments demonstrate that FRI-MxMoE accelerates the profiling phase by up to 15.7× (on DeepSeek-V2) while achieving comparable or slightly superior zero-shot accuracy (e.g., +1.04% on DeepSeekV2-Lite) compared to the baseline. This enables continuous sensitivity modeling, preserves accuracy under mixed-precision allocation, and reduces offline computation by orders of magnitude. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2aea6d0-04af-4554-b319-e27f92730ff3Builds on16
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
Related papers
- MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-DesignHaojie Duanmu, Xiuhong Li, Zhihang Yuan, Size Zheng et al.ICML 2025
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceZhixuan Chen, Xing Hu, Dawei Yang, Zukang Xu et al.ICML 2025
- Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceBaihui Liu, Kaiyuan Tian, Wei Wang, Zhaoning Zhang et al.ACL 2026
- GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMsJianing Deng, Song Wang, Dongwei Wang, Zijie Liu et al.ICML 2026 · 3 citations
- MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware ExpertsWei Tao, Haocheng Lu, Xiaoyang Qu, Bin Zhang et al.ACL 2025 · 8 citations
