EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result Reuse
Fangxin Liu, Ning Yang, Jingkui Yang, Zongwu Wang, Chenyang Guan, Yu Feng, Li Jiang, Haibing Guan
摘要
Mixture-of-Experts (MoE) models significantly reduce computation in large language models by activating only a subset of experts per input token, but they introduce severe memory bottlenecks due to the large number of expert parameters. Existing offloading and prefetching strategies either incur accuracy loss, prohibitively high memory traffic, or high decoding overhead, limiting deployment on resource-constrained hardware. In this work, we present EARTH, a hardware–software co-design that addresses these challenges through three key innovations. First, we propose a dual-entropy encoding scheme that decomposes each expert into a high-information base and a delta component, enabling compact storage while preserving accuracy via adaptive precision management. Second, we introduce a delta-aware speculative prefetching and reuse mechanism that preloads base components of predicted experts and selectively fetches deltas, reusing previously computed delta patterns to reduce memory traffic and redundant computation. Third, we design a hardware accelerator that is co-designed to efficiently support this encoding and prefetching strategy, optimizing execution order, parallelism, and memory utilization. Across representative MoE workloads, EARTH reduces data movement overhead, improves prefetch efficiency, and achieves up to 2.10× speedup compared to state-of-the-art baselines, while maintaining high model accuracy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- SMoE: An Algorithm-System Co-Design for Pushing MoE to the Edge via Expert SubstitutionGuoying Zhu, Meng Li, Haipeng Dai, Xuechen Liu 等ISCA 2026 · 被引用 4 次
- STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE InferenceFangxin Liu, Ning Yang, Zongwu Wang, Chenyang Guan 等ISCA 2026 · 被引用 1 次
- MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingPeng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu 等ASPLOS 2026 · 被引用 4 次
- Self-Speculative Decoding for On-device MoE AccelerationPeirong Zheng, Wenchao Xu, Haozhao WangWWW 2026
- CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesChengcheng Wang, Haowen He, Liang Zhao, Xiaoheng Deng 等AAAI 2026
