Lune

EMNLP2025Top-tier venue

HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference acceleration

Longkai Cheng, Along He, Mulin Li, Xueshuo Xie, Tao Li

2025Year

Abstract

Mixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top-k routing mechanisms. Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant performance degradation. To address this trade-off between efficiency and performance, we propose Hook-MoE, a plug-and-play single-layer compensation framework that effectively restores performance using only a small post-training calibration set. Our method strategically inserts a lightweight trainable Hook module immediately preceding selected transformer blocks. Comprehensive evaluations on four popular MoE models, with an average performance degradation of only 2.5% across various benchmarks, our method reduces the number of activated experts by more than 50% and achieves a 1.42× inference speed-up during the prefill stage. Through systematic analysis, we further reveal that the upper layers require fewer active experts, offering actionable insights for refining dynamic expert selection strategies and enhancing the overall efficiency of MoE models. We make our code available at https://github.com/KerwinKai/HookMoE .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 74aa6c0f-6ad3-4eba-a4b0-a33e7912fc06

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines