McPAL: Scaling Unstructured Sparse Inference with Multi-Chiplet HBM-PIM Architecture for LLMs
Shiwei Liu, Zhirui Huang, Jiangnan Yu, Qi Liu, Chixiao Chen
Abstract
Large language models (LLMs) have gained significant attention recently. However, executing LLM is memory-bound due to the extensive memory accesses. Process-in-memory (PIM) emerges as an energy-efficient solution for LLMs, delivering high memory bandwidth and compute parallelism. Nevertheless, the trend towards larger LLMs introduces escalating memory footprint challenges for monolithic PIM chips. This paper proposes McPAL, which tackles this challenge by emphasizing unstructured sparse compute within PIM and hierarchical multi-chiplet scaling. McPAL decomposes arbitrary sparse weight matrix into multiple irregular sparse vectors. The non-skipped computations in each vector are then routed via an in-memory butterfly network to the standard PIM array, enhancing the PIM utilization. In addition, we scale McPAL vertically by strategically organizing the 3D-HBM hierarchy to minimize the internal long-distance data travel. Meanwhile, a 2.5D IO chiplet scales McPAL horizontally, reducing die-to-die (D2D) data transfer and ensuring sparse workload balance. We conducted extensive experiments from Llama-7B to Llama-70B. The results show that McPAL achieves to speedup and to energy efficiency over Nvidia A100 GPU. Compared to SOTAs, McPAL also achieves to speedup and to energy efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 071fa1b6-c5c2-416e-9270-5b5100bac03aCited by top-tier papers1
Ask how each one uses itRelated papers
- Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationHongyi Li, Songchen Ma, Huanyu Qu, Weihao Zhang et al.ISCA 2026
- MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix PartitionYubin Qin, Yang Wang, Zhiren Zhao, Xiaolong Yang et al.ISCA 2024 · 31 citations
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao et al.ASPLOS 2026 · 1 citation
- PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic LookupYoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim et al.DAC 2025 · 6 citations
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 3 citations
