Lune

ICML2026顶会

Less Token, More Signal: MoE Expert Pruning via Critical Token Selection

Zeliang Zong, Kai Zhang, Yarong Wang, wenming tan, Ye Ren, Jilin Hu

出版方
2026年份

摘要

Mixture-of-Experts (MoE) architectures provide strong scalability for large language models, but their large expert parameter footprint poses challenges for efficient deployment. Expert pruning is widely used to reduce model size and inference cost; however, existing approaches are tokenagnostic, treating all tokens equally when estimating expert importance. This uniform treatment dilutes the contributions of informative tokens and leads to suboptimal pruning decisions. To address this fundamental limitation, we propose STEP (Selective Token-guided Expert Pruning), a token-aware framework that rethinks expert pruning from the perspective of selective token guidance. By incorporating loss-aware expert evaluation and a lightweight knowledge-preserving mechanism, STEP reduces information loss while removing redundant experts. Extensive experiments across different MoE architectures and model scales demonstrate the effectiveness of STEP. On the 30B Qwen3 MoE model with 50% expert sparsity, STEP achieves nearly a 50% reduction in memory usage with minimal performance degradation, delivers a 1.5× throughput improvement and completes the entire pruning process within 10 minutes. Codes are available in https://github.com/hikvision-research/STEP.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖