SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
Rui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu
Abstract
Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) with conditionally-activated parallel experts. However, serving MoE models on memory-constrained devices is challenging due to the large parameter size. Typical solutions such as memory swapping or expert pruning may lead to significantly higher latency or severe accuracy loss. In this paper, we introduce SwapMoE, a framework for efficient serving of MoE-based large language models with tunable memory budgets. The main idea of SwapMoE is to keep a small dynamic set of important experts, namely Virtual Experts, in the main memory for inference, while seamlessly maintaining how the Virtual Experts map to the actual experts. Experiments have shown that SwapMoE can reduce the memory footprint while maintaining reasonable accuracy. For example, on text summarization tasks with Switch Transformer, SwapMoE can reduce the memory consumption from 14.2 GiB to 4.7 GiB, together with 50% latency reduction and a slight Rouge-2 score drop of 0.041.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert ModelsJingcong Liang, Siyuan Wang, Miren Tian, Yitong Li et al.ICLR 2026 · 8 citations
- MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache OptimizationBorui Li, Yitao Wang, Haoran Ma, Ligeng Chen et al.ACL 2025 · 6 citations
- Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained DevicesFahao Chen, Jie Wan, Peng Li, Zhou Su et al.EuroSys 2026 · 2 citations
- CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited MemoryJiashun Suo, Xiaojian Liao, Limin Xiao, Li Ruan et al.ASPLOS 2025 · 1 citation
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
Builds on8
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu et al.MobiCom 2021 · 103 citations
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov et al.EMNLP 2022 · 71 citations
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang et al.MobiCom 2023 · 55 citations
- LegoDNN: block-grained scaling of deep neural networks for mobile visionRui Han, Qinglong Zhang, Chi Harold Liu, Guoren Wang et al.MobiCom 2021 · 51 citations
Related papers
- Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model InferenceJixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao et al.ICML 2025
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime ReconfigurationHamid Reza Imani, Jiaxin Peng, Peiman Mohseni, Abdolah Amirany et al.ICML 2025
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationMengting Ai, Tianxin Wei, Yifan Chen, Zhichen Zeng et al.KDD 2025 · 5 citations
- STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE InferenceFangxin Liu, Ning Yang, Zongwu Wang, Chenyang Guan et al.ISCA 2026 · 1 citation
- MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse ModelsTaehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho et al.DAC 2024 · 11 citations
