Lune

DAC2025顶会

PIMoE: Towards Efficient MoE Transformer Deployment on NPU-PIM System through Throttle-Aware Task Offloading

Lizhou Wu, Haozhe Zhu, Siqi He, Xuanda Lin, Xiaoyang Zeng, Chixiao Chen

2025年份
6被引次数
1顶会引用

摘要

Mixture-of-experts (MoE) technique holds significant promise for scaling up Transformer models. However, the data transfer overhead and imbalanced workload hinder efficient deployment. This work presents PIMoE, a heterogeneous system combining processing-in-memory (PIM) and neural-processing-unit (NPU) to facilitate efficient MoE Transformer inference. We propose a throttle-aware task offloading method that addresses workload imbalance between NPU and PIM, achieving optimal task distribution. Furthermore, we design a near-memory-controller data condenser to address the mismatch of sparse data layout between NPU and PIM, enhancing data transfer efficiency. Experimental results demonstrate that PIMoE achieves 4.5×4.5 \times speedup and 13.7×13.7 \times greater energy efficiency compared to the A 100, and 1.4×1.4 \times speedup over a state-of-the-art MoE platform.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖