Telos: A Dataflow Accelerator for Sparse Triangular Solver of Partial Differential Equations
Xiaochen Hao, Hao Luo, Chu Wang, Chao Yang, Yun Liang
摘要
Partial Differential Equations (PDEs) serve as the backbone of numerous scientific problems.Their solutions often rely on numerical methods, which transform these equations into large, sparse systems of linear equations.These systems, solved with iterative methods, exhibit structured sparsity patterns when derived from stencil-based numerical schemes.In preconditioned solvers, the sparse triangular solve procedure, SpTRSV, usually dominates the entire execution due to its loop-carried dependencies.Optimizing SpTRSV requires extracting parallelism from dependent computations.However, prior works have struggled to achieve both high parallelism and data locality, leading to suboptimal performance.We propose Telos, a dataflow accelerator for SpTRSV that exploits structured sparsity patterns in PDE solving.The dataflow execution leverages stencil patterns, efficiently utilizing pipeline parallelism to resolve data dependencies with minimal overhead.We tackle the challenge of complex data dependencies by proposing a plane-parallel pipelining technique that maps computations onto processing elements (PEs) while preserving data locality.A cross-plane communication aggregation technique is developed to streamline data transfers into a systolic manner.Our accelerator features effective pipelining of dependent computations and overlapping of computations with memory accesses.Experiments demonstrate that Telos delivers average speedups of 61×, 8×, and 11× over CPUs, GPUs, state-of-the-art accelerator, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured GridHuanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu 等OOPSLA 2023 · 被引用 4 次
- Sparsified Preconditioned Conjugate Gradient Solver on GPUsDa Ma, Khalid Ahmad, Kazem Cheshmi, Hari Sundar 等SC 2025 · 被引用 1 次
- Exploiting Computation Reuse for Stencil AcceleratorsYuze Chi, Jason CongDAC 2020 · 被引用 11 次
- Extending Sparse Patterns to Improve Inverse Preconditioning on GPU ArchitecturesSergi Laut, Ricard Borrell, Marc CasasHPDC 2024 · 被引用 3 次
