ReCG: ReRAM-Accelerated Sparse Conjugate Gradient
Mingjia Fan, Xiaoming Chen, Dechuang Yang, Zhou Jin, Weifeng Liu
Abstract
Solving sparse linear systems is crucial in scientific computing. Sparse Conjugate Gradient (CG) is one of the most well-known iterative solvers with high efficiency and low storage requirements. However, the performance of sparse CG solvers implemented on storage-compute separated architectures is greatly limited by the irregular memory access and the large amount of data transmission.
In this paper, we propose a processing-in-memory (PIM) architecture, ReCG, based on the resistive random-access memory (ReRAM) to accelerate sparse CG solvers. The design of ReCG faces three major challenges: (1) how to make complex sparse CG more suitable for acceleration with ReRAM-based architecture, (2) how to map sparse and irregular operations to regular crossbars that are more suitable for dense operations, and (3) how to coordinate the dataflow among hardware units to minimize the impact of the poor write endurance of ReRAMs on CG acceleration. To address these challenges, we (1) classify the kernels of sparse CG by exploring the commonality of operations and design a flexible and dedicated architecture, (2) efficiently implement the sparse and irregular operations by utilizing both content-addressable memory (CAM) and multiply-and-accumulate (MAC) crossbars, and (3) develop a novel scheduling strategy for the dataflow. The experimental results show that ReCG improves the performance by up to three, one and one order of magnitude compared with PETSc on CPU and GPU and CALLIPEPLA on FPGA, respectively, and the energy consumption is reduced by up to two, two and one order of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 708ad531-8569-40c6-ba8d-d8ef51595a54Cited by top-tier papers2
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu et al.SC 2024 · 17 citations
- Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsDechuang Yang, Yuxuan Zhao, Yiduo Niu, Weile Jia et al.SC 2024 · 8 citations
Builds on4
- GaaS-X: Graph Analytics Accelerator Supporting Sparse Data Representation using Crossbar ArchitecturesNagadastagiri Challapalle, Sahithi Rampalli, Linghao Song, Nandhini Chandramoorthy et al.ISCA 2020 · 67 citations
- PIM-DH: ReRAM-based processing-in-memory architecture for deep hashing accelerationFangxin Liu, Wenbo Zhao, Yongbiao Chen, Zongwu Wang et al.DAC 2022 · 11 citations
- FSPA: An FeFET-based Sparse Matrix-Dense Vector Multiplication AcceleratorXiaoyu Zhang, Zerun Li, Rui Liu, Xiaoming Chen et al.DAC 2023 · 10 citations
- AmgR: Algebraic Multigrid Accelerated on ReRAMMingjia Fan, Xiaotian Tian, Yintao He, Junxian Li et al.DAC 2023 · 8 citations
Related papers
- ReSMiPS: A ReRAM-based Sparse Mixed-precision Solver with Fast Matrix Reordering AlgorithmYuyang Fu, Jiancong Li, Jia Chen, Zhiwei Zhou et al.DAC 2025
- PIMGCN: A ReRAM-Based PIM Design for Graph Convolutional Network AccelerationTao Yang, Dongyue Li, Yibo Han, Yilong Zhao et al.DAC 2021 · 39 citations
- ReSiPE: ReRAM-based Single-Spiking Processing-In-Memory EngineZiru Li, Bonan Yan, Hai Helen LiDAC 2020 · 18 citations
- ReFloat: Low-Cost Floating-Point Processing in ReRAM for Accelerating Iterative Linear SolversLinghao Song, Fan Chen, Hai Li, Yiran ChenSC 2023 · 8 citations
- ReSMA: accelerating approximate string matching using ReRAM-based content addressable memoryHuize Li, Hai Jin, Long Zheng, Yu Huang et al.DAC 2022 · 7 citations
