DIAMoND: Dynamic Inference for Adaptive Edge MOE with Heterogeneous In-NAND and Near-DRAM Compute Architecture
Ling Liang, Tianyang Luo, Shuzhang Zhong, Dongxue Zhao, Qichao Ma, Renjie Wei, Jingyu Wang, Meng Li, Guangyu Sun, Zongwei Wang, Yimao Cai
Abstract
Among modern large language models (LLMs), the Mixture-of-Experts (MoE) model stands out as a promising approach. Although MoE models activate only small portions of experts during inference, the full model size can reach 50 100GB, posing challenges for the memory capacity of edge devices. Additionally, single-batch decoding in edge applications places significant memory bandwidth demands for parameter loading, and the dynamic expert selection in MoE models further complicates loading patterns. Previous studies have explored using NAND-Flash-based SSD to provide the large storage capacity needed for holding LLM parameters. However, the performance of near-NAND computing remains limited by the memory data transfer rate, and these studies lack designs to accommodate the complex execution flow of MoE models. To fully release the computational capacity on the memory side while considering memory characteristics, we introduce DIAMoND, a heterogeneous accelerator that integrates in-NAND and near-DRAM computing via a 2.5D package to support all operations in MoE inference efficiently. To address mismatches between varying matrix sizes and the fixed NAND array size, we propose a mask-based mapping method under in-NAND computing. Finally, a dynamic online expert selection scheme based on in-NAND computing is proposed to enhance MoE inference efficiency. Overall, our proposed architecture enables the edge inference of Mixtral-8×7B at a speed of 197.3 tokens/s and a peak energy efficiency of 5.8 tokens/J, with the speed being and better than GPU and ASIC LLM edge accelerators.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0a6b5fe8-4987-4fa2-bda5-45ed475a2f04Related papers
- MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUsShiyi Cao, Shu Liu, Tyler Griggs, Peter Schafhalter et al.ASPLOS 2025 · 15 citations
- MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse ModelsTaehyun Kim, Kwanseok Choi, Youngmock Cho, Jaehoon Cho et al.DAC 2024 · 11 citations
- Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model InferenceJixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao et al.ICML 2025
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu et al.ISCA 2026 · 4 citations
- Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic ParallelizationJiahao Chen, Shigang Li, Rongtian Fu, Tong Wu et al.HPDC 2026
