STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE Inference
Fangxin Liu, Ning Yang, Zongwu Wang, Chenyang Guan, Haomin Li, Yu Feng, Liqiang Lu, Xiang Li, Siran Yang, Jiamang Wang, Lin Qu, Li Jiang, Haibing Guan
摘要
Large language models (LLMs) have driven significant advances in natural language processing, but their substantial memory and computational demands pose challenges for real-time deployment in memory- and bandwidth-constrained systems. Sparse Mixture-of-Experts (MoE) models mitigate these issues by selectively activating a subset of parameters per input, yet their irregular memory access patterns introduce significant latency and memory overhead. We propose STEP, an adaptive inference framework that optimizes MoE inference through spatio-temporal expert prefetching. STEP introduces layer-wise expert allocation to dynamically adjust the number of activated experts based on computational importance, reducing unnecessary computation and memory traffic. Additionally, it employs a predictive expert prefetching mechanism that leverages temporal locality and layer-specific predictability patterns to minimize access latency. By incorporating a token-aware adaptive window mechanism, STEP further enhances prefetch accuracy. Extensive evaluations on representative MoE models demonstrate that STEP achieves up to speedup while maintaining model accuracy, outperforming state-of-the-art baselines in latency and memory efficiency under constrained memory budgets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesChengcheng Wang, Haowen He, Liang Zhao, Xiaoheng Deng 等AAAI 2026
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu 等ASPLOS 2026 · 被引用 4 次
- Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceBaihui Liu, Kaiyuan Tian, Wei Wang, Zhaoning Zhang 等ACL 2026
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng 等SC 2025 · 被引用 3 次
- FIRM-MoE: Fine-GrainedExpert Decomposition for Resource-Adaptive MoE InferenceKeyu Chen, Qihang Zhou, Bin Qian, Zhenyu Wen 等AAAI 2026
