Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference
Jixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao, Mengyi Chen, Yifeng Yang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Dongsheng Li, David A. Clifton, Qin Lv
Abstract
Mixture-of-Experts (MoE) is widely adopted to deploy Large Language Models (LLMs) on edge devices with limited memory budgets. Although MoE is, in theory, an inborn memory-friendly architecture requiring only a few activated experts to reside in the memory for inference, current MoE architectures cannot effectively fulfill this advantage and will yield intolerable inference latencies of LLMs on memory-constrained devices. Our investigation pinpoints the essential cause as the remarkable temporal inconsistencies of inter-token expert activations, which generate overly frequent expert swapping demands dominating the latencies. To this end, we propose a novel MoE architecture, Oracle-MoE, to fulfill the real on-device potential of MoE-based LLMs. Oracle-MoE route tokens in a highly compact space suggested by attention scores, termed the oracle space, to effectively maintain the semantic locality across consecutive tokens to reduce expert activation variations, eliminating massive swapping demands. Theoretical analysis proves that Oracle-MoE is bound to provide routing decisions with better semantic locality and, there-* Equal contribution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cea6ee82-1362-412f-91c4-bfee806e990eCited by top-tier papers3
- ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM InferenceXiongwei Zhu, Xiaojian Liao, Tianyang Jiang, Yusen Zhang et al.ICML 2026 · 2 citations
- SD-MoE: Spectral Decomposition for Effective Expert SpecializationRuijun Huang, Fang DONG(董方), Xin Zhang, Anrui Chen et al.ICML 2026 · 2 citations
- Building Massively Multimodal Foundation Models with Interaction-aware Mixture-of-ExpertsXing Han, Hsing-Huan Chung, Joydeep Ghosh, Paul Pu Liang et al.ICLR 2026 · 1 citation
Builds on12
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
- Unified Scaling Laws for Routed Language ModelsAidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch et al.ICML 2022 · 266 citations
- On the Representation Collapse of Sparse Mixture of ExpertsZewen Chi, Li Dong, Shaohan Huang, Damai Dai et al.NeurIPS 2022 · 223 citations
Related papers
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu et al.ASPLOS 2026 · 4 citations
- FIRM-MoE: Fine-GrainedExpert Decomposition for Resource-Adaptive MoE InferenceKeyu Chen, Qihang Zhou, Bin Qian, Zhenyu Wen et al.AAAI 2026
- ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity SchedulingYuchen Yang, Yaru Zhao, Pu Yang, Shaowei Wang et al.ICML 2026 · 3 citations
- SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory BudgetRui Kong, Yuanchun Li, Qingtian Feng, Weijun Wang et al.ACL 2024 · 12 citations
- Fate: Fasss sEsdge Inference of Mixture-of-Experts Models via Cross-Layer GateZhiyuan Fang, Xingfan Yu, Yuegui Huang, Zicong Hong et al.WWW 2026 · 4 citations
