Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference
Jixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao, Mengyi Chen, Yifeng Yang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Dongsheng Li, David A. Clifton, Qin Lv
摘要
Mixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures cannot effectively fulfill this advantage and will yield intolerable inference latencies of LLMs on memory-constrained devices. Our investigation pinpoints the essential cause as the remarkable temporal inconsistencies of inter-token expert activations, which generate overly frequent expert swapping demands dominating the latencies. To this end, we propose a novel MoE architecture, Oracle-MoE, to fulfill the real on-device potential of MoE-based LLMs. Oracle-MoE route tokens in a highly compact space suggested by attention scores, termed the oracle space, to effectively maintain the semantic locality across consecutive tokens to reduce expert activation variations, eliminating massive swapping demands. Theoretical analysis proves that Oracle-MoE is bound to provide routing decisions with better semantic locality and, there-* Equal contribution
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM InferenceXiongwei Zhu, Xiaojian Liao, Tianyang Jiang, Yusen Zhang 等ICML 2026 · 被引用 2 次
- SD-MoE: Spectral Decomposition for Effective Expert SpecializationRuijun Huang, Fang DONG(董方), Xin Zhang, Anrui Chen 等ICML 2026 · 被引用 2 次
- Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-ExpertsXing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper12
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch 等ICML 2022 · 被引用 266 次
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai 等NeurIPS 2022 · 被引用 223 次
相关 Paper
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu 等ASPLOS 2026 · 被引用 4 次
- FIRM-MoE: Fine-GrainedExpert Decomposition for Resource-Adaptive MoE InferenceKeyu Chen, Qihang Zhou, Bin Qian, Zhenyu Wen 等AAAI 2026
- ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity SchedulingYuchen Yang, Yaru Zhao, Pu Yang, Shaowei Wang 等ICML 2026 · 被引用 3 次
- SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang 等ACL 2024 · 被引用 12 次
- Fate: Fasss sEsdge Inference of Mixture-of-Experts Models via Cross-Layer GateZhiyuan Fang, Xingfan Yu, Yuegui Huang, Zicong Hong 等WWW 2026 · 被引用 4 次
