Lune

ACL2026顶会

Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation

Huachen Qi, Ruiyu Zhuo, Bowen Shi, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen

2026年份

摘要

Large Language Models continue to scale in size and capability, driving substantial computational and memory demands. Mixtureof-Experts (MoE) architectures alleviate this cost by activating only a sparse subset of experts per token, enabling efficient scaling without proportional increases in inference compute. However, quantization in MoE models remains challenging due to heterogeneous sensitivity across experts and their internal linear layers. Existing mixed-precision frameworks such as Mixed-precision Quantization for MoE (MxMoE) require full quantization-loss evaluation for expert-layer-and-bit configurations, incurring prohibitive profiling cost. To address this, we propose FRI-MxMoE, a profilingfree mixed-precision quantization framework that reformulates MoE calibration from exhaustive expert-wise profiling to sparse anchor profiling followed by Fuzzy Rule Interpolation. By constructing a fuzzy rule base in the intraexpert layer feature space (bit-width, activation variance, parameter scale), our method predicts quantization error from only sparse samples while remaining compatible with existing mixed-precision allocation objectives. Extensive experiments demonstrate that FRI-MxMoE accelerates the profiling phase by up to 15.7× (on DeepSeek-V2) while achieving comparable or slightly superior zero-shot accuracy (e.g., +1.04% on DeepSeekV2-Lite) compared to the baseline. This enables continuous sensitivity modeling, preserves accuracy under mixed-precision allocation, and reduces offline computation by orders of magnitude. 1

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext e2aea6d0-04af-4554-b319-e27f92730ff3

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖