Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation
Huachen Qi, Ruiyu Zhuo, Bowen Shi, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen
摘要
Large Language Models continue to scale in size and capability, driving substantial computational and memory demands. Mixtureof-Experts (MoE) architectures alleviate this cost by activating only a sparse subset of experts per token, enabling efficient scaling without proportional increases in inference compute. However, quantization in MoE models remains challenging due to heterogeneous sensitivity across experts and their internal linear layers. Existing mixed-precision frameworks such as Mixed-precision Quantization for MoE (MxMoE) require full quantization-loss evaluation for expert-layer-and-bit configurations, incurring prohibitive profiling cost. To address this, we propose FRI-MxMoE, a profilingfree mixed-precision quantization framework that reformulates MoE calibration from exhaustive expert-wise profiling to sparse anchor profiling followed by Fuzzy Rule Interpolation. By constructing a fuzzy rule base in the intraexpert layer feature space (bit-width, activation variance, parameter scale), our method predicts quantization error from only sparse samples while remaining compatible with existing mixed-precision allocation objectives. Extensive experiments demonstrate that FRI-MxMoE accelerates the profiling phase by up to 15.7× (on DeepSeek-V2) while achieving comparable or slightly superior zero-shot accuracy (e.g., +1.04% on DeepSeekV2-Lite) compared to the baseline. This enables continuous sensitivity modeling, preserves accuracy under mixed-precision allocation, and reduces offline computation by orders of magnitude. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-DesignHaojie Duanmu, Xiuhong Li, Zhihang Yuan, Size Zheng 等ICML 2025
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceZhixuan Chen, Xing Hu, Dawei Yang, Zukang Xu 等ICML 2025
- Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceBaihui Liu, Kaiyuan Tian, Wei Wang, Zhaoning Zhang 等ACL 2026
- GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMsJianing Deng, Song Wang, Dongwei Wang, Zijie Liu 等ICML 2026 · 被引用 3 次
- MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware ExpertsWei Tao, Haocheng Lu, Xiaoyang Qu, Bin Zhang 等ACL 2025 · 被引用 8 次
