Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert Caching
Kexin Li, Wenkan Huang, Qinggang Wang, Long Zheng, Xiaofei Liao, Hai Jin, Jingling Xue
摘要
The emerging Mixture-of-Experts (MoE) model mitigates the high compute cost of large-scale LLMs by sparsely activating a subset of experts during inference. However, MoE requires storing massive expert parameters, creating a severe memory bottleneck on resource-constrained GPUs. Existing approaches offload parameters to host memory and prefetch activated experts to GPU memory with sophisticated policies, but these solutions are tailored to single-batch inference and suffer from communication bottlenecks at larger batch sizes, limiting throughput. We identify two forms of locality in expert activation: a small set of experts are frequently invoked across inference (global locality), while others recur within short decoding bursts (temporal locality). To exploit this, we propose Diff-MoE, which introduces a differential cache hierarchy in GPU memory. Globally hot experts reside in per-layer high-priority caches, locally hot ones are dynamically managed in per-layer medium-priority caches under a priority-driven replacement policy, and the remaining cold experts are cached temporarily and evicted on demand. Moreover, Diff-MoE incorporates a lightweight predictor that prefetches experts likely needed in the next MoE layer, overlapping migration with computation to further reduce latency. Our evaluation shows that Diff-MoE improves inference throughput by 2.74 ×, 2.22 ×, and 1.55 × over DeepSpeed, Pre-gated MoE, and MoE-Infinity, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse ModelsTaehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho 等DAC 2024 · 被引用 11 次
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu 等ISCA 2026 · 被引用 4 次
- Toward Efficient Inference for Mixture of ExpertsHaiyang Huang, Newsha Ardalani, Anna Y. Sun, Liu Ke 等NeurIPS 2024 · 被引用 60 次
- STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE InferenceFangxin Liu, Ning Yang, Zongwu Wang, Chenyang Guan 等ISCA 2026 · 被引用 1 次
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu 等ASPLOS 2026 · 被引用 4 次
