AmgR: Algebraic Multigrid Accelerated on ReRAM
Mingjia Fan, Xiaotian Tian, Yintao He, Junxian Li, Yiru Duan, Xiaozhe Hu, Ying Wang, Zhou Jin, Weifeng Liu
Abstract
Solving systems of linear equations is a fundamental problem in scientific computing, which has been extensively researched for decades. One of the most well-known solvers is Algebraic Multigrid (AMG), which is widely used in high performance computing due to its good scalability. But currently accelerating AMG relies on the traditional von Neumann architecture of storage and computation separation, which leads to a large data transmission overhead. In this work, we propose a ReRAM-based processing-in-memory (PIM) architecture named AmgR, which overcomes the limitations of the traditional von Neumann architecture for AMG acceleration.
However, accelerating AMG on ReRAM is non-trivial, because (1) AMG has many computing kernels of various types; (2) there are irregular operations that cannot be directly performed using matrix-vector multiplication suitable for ReRAM, i.e., aggregation operation; (3) ReRAM has poor write endurance, and a lot of data during AMG acceleration needs to be rewritten into ReRAM, resulting in high write cost. To address these issues, firstly, we propose a flexible architecture, which can realize each kernel of AMG and is reused by many kernels to improve resource utilization. Secondly, we propose a dedicated unit to realize the aggregation operation. Finally, we present a new mapping strategy to greatly reduce the number of data handling and writes. The experimental results show that the performance of AmgR is improved by an average of one and two orders of magnitude compared to HYPRE on the CPU and AmgX on the GPU, respectively, while the energy consumption is reduced by an average of two and three orders of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75466e59-09d6-4e3c-ba45-f0840cb7f619Cited by top-tier papers3
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu et al.SC 2024 · 17 citations
- Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsDechuang Yang, Yuxuan Zhao, Yiduo Niu, Weile Jia et al.SC 2024 · 8 citations
- ReCG: ReRAM-Accelerated Sparse Conjugate GradientMingjia Fan, Xiaoming Chen, Dechuang Yang, Zhou Jin et al.DAC 2024 · 5 citations
Builds on1
Related papers
- ReGNN: a ReRAM-based heterogeneous architecture for general graph neural networksCong Liu, Haikun Liu, Hai Jin, Xiaofei Liao et al.DAC 2022 · 22 citations
- ReSMA: accelerating approximate string matching using ReRAM-based content addressable memoryHuize Li, Hai Jin, Long Zheng, Yu Huang et al.DAC 2022 · 7 citations
- TARe: Task-Adaptive in-situ ReRAM Computing for Graph LearningYintao He, Ying Wang, Cheng Liu, Huawei Li et al.DAC 2021 · 12 citations
- ReFloat: Low-Cost Floating-Point Processing in ReRAM for Accelerating Iterative Linear SolversLinghao Song, Fan Chen, Hai Li, Yiran ChenSC 2023 · 8 citations
- Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationYue Zha, Jing LiISCA 2020 · 31 citations
