pSyncPIM: Partially Synchronous Execution of Sparse Matrix Operations for All-Bank PIM Architectures
Daehyeon Baek, Soojin Hwang, Jaehyuk Huh
摘要
Recent commercial incarnations of processing-in-memory (PIM) maintain the standard DRAM interface and employ the all-bank mode execution to maximize bank-level memory bandwidth. Such a synchronized all-bank PIM control can effectively manage conventional dense matrix-vector operations on evenly distributed matrices across banks with lock-step execution. Sparse matrix processing is another critical computation that can significantly benefit from the PIM architecture, but the current all-bank PIM control cannot support diverging executions due to the random sparsity. To accelerate such sparse matrix applications, this paper proposes a partially synchronous execution on sparse matrix-vector multiplication (SpMV) and sparse triangular matrix-vector solve (SpTRSV), filling the gap between the practical constraint of PIM and the irregular nature of sparse computation. It allows the execution of the processing unit of each bank to diverge in a limited way to manage the irregular execution path of sparse matrix computation. It proposes compaction and distribution policies for the input matrix and vector. In addition to SpMV, this paper identifies SpTRSV is another key kernel, and proposes SpTRSV acceleration on PIM technology. The experimental evaluation shows that the new sparse PIM architecture outperforms NVIDIA Geforce RTX 3080 GPU by speedup for SpMV and speedup for SpTRSV with a similar amount of DRAM bandwidth.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model ServingWonung Kim, Yubin Lee, Yoonsung Kim, Jinwoo Hwang 等MICRO 2025 · 被引用 6 次
- PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage FormatYilong Zhao, Mingyu Gao, Huanchen Zhang, Fangxin Liu 等ASPLOS 2025 · 被引用 4 次
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao 等ISCA 2025 · 被引用 1 次
- PIM-Malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) ArchitecturesDongjae Lee, Bongjoon Hyun, Youngjin Kwon, Minsoo RhuHPCA 2026 · 被引用 1 次
- LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMJunguk Hong, Changmin Shin, Sukjin Kim, Si Ung Noh 等HPCA 2026
相关 Paper
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak 等HPCA 2021 · 被引用 111 次
- Gearbox: a case for supporting accumulation dispatching and hybrid partitioning in PIM-based acceleratorsMarzieh Lenjani, Alif Ahmed, Mircea Stan, Kevin SkadronISCA 2022 · 被引用 21 次
- AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryHongju Kal, Chanyoung Yoo, Won Woo RoMICRO 2023 · 被引用 19 次
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 被引用 37 次
