Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
Chenpeng Wu, Qiqi Gu, Heng Shi, Jianguo Yao, Haibing Guan
Abstract
The escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to enhance efficiency without compromising model accuracy. Structured sparsity emerges as a compelling strategy to address these challenges by leveraging the emerging sparse computing hardware. Prior works mainly focus on the sparsity in model parameters, neglecting the inherent sparse patterns in activations. This oversight can lead to additional computational costs associated with activations, potentially resulting in suboptimal performance.
This paper presents Samoyeds, an innovative acceleration system for MoE LLMs utilizing Sparse Tensor Cores (SpTCs). Samoyeds is the first to apply sparsity simultaneously to both activations and model parameters. It introduces a bespoke sparse data format tailored for MoE computation and develops a specialized sparse-sparse matrix multiplication kernel. Furthermore, Samoyeds incorporates systematic optimizations specifically designed for the execution of dual-side structured sparse MoE LLMs on SpTCs, further enhancing system performance. Evaluations show that Samoyeds outperforms SOTA works by up to 1.99× at the kernel level and 1.58× at the model level. Moreover, it enhances memory
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b32047df-3450-41fa-85b5-303bb4df6017Cited by top-tier papers4
- SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingQiqi Gu, Chenpeng Wu, Heng Shi, Jianguo YaoPPoPP 2026 · 1 citation
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
- High Throughput and Low Latency LLM Serving via Adaptive KV CachingWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2026
- TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive SchedulingJunyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao et al.EuroSys 2026
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource ConstraintsYichao Yuan, Lin Ma, Nishil TalatiHPDC 2026
- Masks Can be Learned as an Alternative to ExpertsPeiyu Liu, Tianwen Wei, Bo Zhu, Xin Zhao et al.ACL 2025 · 1 citation
- SonicMoE: Accelerating MoE with IO and Tile-aware OptimizationsWentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica et al.ICLR 2026 · 22 citations
- Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferenceRanggi Hwang, Jianyu Wei, Shijie Cao, Changho Hwang et al.ISCA 2024 · 48 citations
- Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and InferenceWeilin Cai, Le Qin, Shwai He, Junwei Cui et al.ICML 2026
