AmgR: Algebraic Multigrid Accelerated on ReRAM
Mingjia Fan, Xiaotian Tian, Yintao He, Junxian Li, Yiru Duan, Xiaozhe Hu, Ying Wang, Zhou Jin, Weifeng Liu
摘要
Solving systems of linear equations is a fundamental problem in scientific computing, which has been extensively researched for decades. One of the most well-known solvers is Algebraic Multigrid (AMG), which is widely used in high performance computing due to its good scalability. But currently accelerating AMG relies on the traditional von Neumann architecture of storage and computation separation, which leads to a large data transmission overhead. In this work, we propose a ReRAM-based processing-in-memory (PIM) architecture named AmgR, which overcomes the limitations of the traditional von Neumann architecture for AMG acceleration.
However, accelerating AMG on ReRAM is non-trivial, because (1) AMG has many computing kernels of various types; (2) there are irregular operations that cannot be directly performed using matrix-vector multiplication suitable for ReRAM, i.e., aggregation operation; (3) ReRAM has poor write endurance, and a lot of data during AMG acceleration needs to be rewritten into ReRAM, resulting in high write cost. To address these issues, firstly, we propose a flexible architecture, which can realize each kernel of AMG and is reused by many kernels to improve resource utilization. Secondly, we propose a dedicated unit to realize the aggregation operation. Finally, we present a new mapping strategy to greatly reduce the number of data handling and writes. The experimental results show that the performance of AmgR is improved by an average of one and two orders of magnitude compared to HYPRE on the CPU and AmgX on the GPU, respectively, while the energy consumption is reduced by an average of two and three orders of magnitude.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu 等SC 2024 · 被引用 17 次
- Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsDechuang Yang, Yuxuan Zhao, Yiduo Niu, Weile Jia 等SC 2024 · 被引用 8 次
- ReCG: ReRAM-Accelerated Sparse Conjugate GradientMingjia Fan, Xiaoming Chen, Dechuang Yang, Zhou Jin 等DAC 2024 · 被引用 5 次
它引用的顶会 Paper1
相关 Paper
- ReGNN: a ReRAM-based heterogeneous architecture for general graph neural networksCong Liu, Haikun Liu, Hai Jin, Xiaofei Liao 等DAC 2022 · 被引用 22 次
- ReSMA: accelerating approximate string matching using ReRAM-based content addressable memoryHuize Li, Hai Jin, Long Zheng, Yu Huang 等DAC 2022 · 被引用 7 次
- TARe: Task-Adaptive in-situ ReRAM Computing for Graph LearningYintao He, Ying Wang, Cheng Liu, Huawei Li 等DAC 2021 · 被引用 12 次
- ReFloat: Low-Cost Floating-Point Processing in ReRAM for Accelerating Iterative Linear SolversLinghao Song, Fan Chen, Hai Li, Yiran ChenSC 2023 · 被引用 8 次
- Hyper-Ap: Enhancing Associative Processing Through A Full-Stack OptimizationYue Zha, Jing LiISCA 2020 · 被引用 31 次
