CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory Constraints
Han Li, Jingwei Sun, Junqing Lin, Guangzhong Sun
Abstract
Mixture of Experts (MoE) models have emerged as a promising approach to scale language models efficiently by activating only a subset of parameters for each input. However, deploying these models under GPU memory constraints remains challenging, as existing offloading strategies incur significant overhead from CPU-GPU data transfers. While prior work has explored prefetching techniques to mitigate this bottleneck, these methods require costly fallback mechanisms when predictions fail. Since expert transfers cannot be canceled once initiated, the correct experts need to be loaded on demand sequentially, introducing additional latency. To address this, we present CommitMoE, a novel approach featuring a Commit Router that makes execution decisions based on expert predictions without fallback mechanisms. Our key insight reveals that router certainty strongly correlates with prediction accuracy, while in low-certainty scenarios, the model output demonstrates inherent robustness to expert selection. Leveraging this insight to design a systems-level solution, CommitMoE achieves 1.3× to 9.4× faster inference across different environments and datasets compared to state-of-the-art offloading frameworks while maintaining model quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7915a13-3a4d-4b5f-926b-aeafd7677657Builds on13
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferenceRanggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang et al.ISCA 2024 · 48 citations
Related papers
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu et al.ISCA 2026 · 4 citations
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert ModelsJingcong Liang, Siyuan Wang, Miren Tian, Yitong Li et al.ICLR 2026 · 8 citations
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
- FIRM-MoE: Fine-GrainedExpert Decomposition for Resource-Adaptive MoE InferenceKeyu Chen, Qihang Zhou, Bin Qian, Zhenyu Wen et al.AAAI 2026
- CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesChengcheng Wang, Haowen He, Liang Zhao, Xiaoheng Deng et al.AAAI 2026
