SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator
Xinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak, Lei Deng, Ling Liang, Xing Hu, Yuan Xie
Abstract
Sparse matrix-vector multiplication (SpMV) is an important primitive across a wide range of application domains such as scientific computing and graph analytics. Due to its intrinsic memory-bound characteristics, the performance of SpMV on throughput-oriented architectures such as GPU is bounded by the limited bandwidth between processors and memory. Processing-in-memory (PIM) architectures, made feasible by advances in 3D stacking, provide new opportunities to utilize ultra-high bandwidth by integrating compute-logic into memory.In this paper, we develop an SpMV accelerator, named as SpaceA, based on PIM architectures. SpaceA integrates compute logic near memory banks to exploit bank-level bandwidth. SpaceA contains both hardware and data-mapping design features to alleviate irregular memory access patterns which hinder full utilization of high memory bandwidth. In terms of hardware design features, SpaceA consists of two unique features: (1) it utilizes the capability of outstanding memory requests to hide the memory access latency to data located in non-local memory banks; (2) it integrates Content Addressable Memory (CAM) at the bank level to exploit data reuse of the input vectors. In addition, we develop a mapping scheme that partitions the sparse matrix into different memory banks, to maximize the data locality of the input vector and to achieve workload balance among processing elements (PEs) near each bank. Overall, SpaceA together with the proposed mapping method achieves 13.54x speedup and 87.49% energy saving on average over the GPU baseline on SpMV computation. In addition to SpMV primitives, we conduct a case study on graph analytics to demonstrate the benefits of SpaceA for applications built on SpMV. Compared to Tesseract and GraphP, state-of-the-art graph accelerators, SpaceA obtains better performance due to its higher effective bandwidth provided by near-bank integration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- Serpens: a high bandwidth memory based accelerator for general-purpose sparse matrix-vector multiplicationLinghao Song, Yuze Chi, Licheng Guo, Jason CongDAC 2022 · 56 citations
- PIM-tree: A Skew-resistant Index for Processing-in-MemoryHongbo Kang, Yiwei Zhao, Guy E. Blelloch, Laxman Dhulipala et al.VLDB 2023 · 40 citations
- ABNDP: Co-optimizing Data Access and Load Balance in Near-Data ProcessingBoyu Tian, Qihang Chen, Mingyu GaoASPLOS 2023 · 31 citations
- MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflowsSiying Feng, Xin He, Kuan-Yu Chen, Liu Ke et al.ISCA 2022 · 28 citations
Builds on5
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank ArchitecturePeng Gu, Xinfeng Xie, Yufei Ding, Guoyang Chen et al.ISCA 2020 · 77 citations
- ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorBahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim et al.HPCA 2020 · 63 citations
- DUET: Boosting Deep Neural Network Efficiency on Dual-Module ArchitectureLiu Liu, Zheng Qu, Lei Deng, Fengbin Tu et al.MICRO 2020 · 27 citations
Related papers
- pSyncPIM: Partially Synchronous Execution of Sparse Matrix Operations for All-Bank PIM ArchitecturesDaehyeon Baek, Soojin Hwang, Jaehyuk HuhISCA 2024 · 18 citations
- GaaS-X: Graph Analytics Accelerator Supporting Sparse Data Representation using Crossbar ArchitecturesNagadastagiri Challapalle, Sahithi Rampalli, Linghao Song, Nandhini Chandramoorthy et al.ISCA 2020 · 67 citations
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 37 citations
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi et al.MICRO 2020 · 223 citations
- 3D-SubG: A 3D Stacked Hybrid Processing Near/In-Memory Accelerator for Subgraph GNNsGuoxiang Li, Runnan Xu, Ruohang Xu, Yikan Qiu et al.DAC 2025 · 2 citations
