SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
Rui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang, Xiaozhou Ye, Ye Ouyang, Linghe Kong, Yunxin Liu
摘要
Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) with conditionally-activated parallel experts. However, serving MoE models on memory-constrained devices is challenging due to the large parameter size. Typical solutions such as memory swapping or expert pruning may lead to significantly higher latency or severe accuracy loss. In this paper, we introduce SwapMoE, a framework for efficient serving of MoE-based large language models with tunable memory budgets. The main idea of SwapMoE is to keep a small dynamic set of important experts, namely Virtual Experts, in the main memory for inference, while seamlessly maintaining how the Virtual Experts map to the actual experts. Experiments have shown that SwapMoE can reduce the memory footprint while maintaining reasonable accuracy. For example, on text summarization tasks with Switch Transformer, SwapMoE can reduce the memory consumption from 14.2 GiB to 4.7 GiB, together with 50% latency reduction and a slight Rouge-2 score drop of 0.041.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert ModelsJingcong Liang, Siyuan Wang, Miren Tian, Yitong Li 等ICLR 2026 · 被引用 8 次
- MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache OptimizationBorui Li, Yitao Wang, Haoran Ma, Ligeng Chen 等ACL 2025 · 被引用 6 次
- Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained DevicesFahao Chen, Jie Wan, Peng Li, Zhou Su 等EuroSys 2026 · 被引用 2 次
- CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited MemoryJiashun Suo, Xiaojian Liao, Limin Xiao, Li Ruan 等ASPLOS 2025 · 被引用 1 次
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang 等EuroSys 2026 · 被引用 1 次
它引用的顶会 Paper8
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Flexible high-resolution object detection on edge devices with tunable latencyShiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu 等MobiCom 2021 · 被引用 103 次
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov 等EMNLP 2022 · 被引用 71 次
- AdaptiveNet: Post-deployment Neural Architecture Adaptation for Diverse Edge EnvironmentsHao Wen, Yuanchun Li, Zunshuai Zhang, Shiqi Jiang 等MobiCom 2023 · 被引用 55 次
- LegoDNN: block-grained scaling of deep neural networks for mobile visionRui Han, Qinglong Zhang, Chi Harold Liu, Guoren Wang 等MobiCom 2021 · 被引用 51 次
相关 Paper
- Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model InferenceJixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao 等ICML 2025
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime ReconfigurationHamid Reza Imani, Jiaxin Peng, Peiman Mohseni, Abdolah Amirany 等ICML 2025
- ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationMengting Ai, Tianxin Wei, Yifan Chen, Zhichen Zeng 等KDD 2025 · 被引用 5 次
- STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE InferenceFangxin Liu, Ning Yang, Zongwu Wang, Chenyang Guan 等ISCA 2026 · 被引用 1 次
- MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse ModelsTaehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho 等DAC 2024 · 被引用 11 次
