Utilizing Sparsity in the GPU-accelerated Assembly of Schur Complement Matrices in Domain Decomposition Methods
Jakub Homola, Ondrej Meca, Lubomír Ríha, Tomás Brzobohatý
摘要
Schur complement matrices emerge in many domain decomposition methods that can solve complex engineering problems using supercomputers. Today, as most of the high-performance clusters' performance lies in GPUs, these methods should also be accelerated.
Typically, the offloaded components are the explicitly assembled dense Schur complement matrices used later in the iterative solver for multiplication with a vector. As the explicit assembly is expensive, it represents a significant overhead associated with this approach to acceleration. It has already been shown that the overhead can be minimized by assembling the Schur complements directly on the GPU.
This paper shows that the GPU assembly can be further improved by wisely utilizing the sparsity of the input matrices. In the context of FETI methods, we achieved a speedup of 5.1 in the GPU section of the code and 3.3 for the whole assembly, making the acceleration beneficial from as few as 10 iterations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Solving Linear Systems on a GPU with Hierarchically Off-Diagonal Low-Rank ApproximationsChao Chen, Per-Gunnar MartinssonSC 2022 · 被引用 4 次
- Second-order Stencil Descent for Interior-point HyperelasticityLei Lan, Minchen Li, Chenfanfu Jiang, Huamin Wang 等SIGGRAPH 2023 · 被引用 33 次
- Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct SolversAhmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov 等SC 2022 · 被引用 4 次
- A GPU-based multilevel additive schwarz preconditioner for cloth and deformable body simulationBotao Wu, Zhendong Wang, Huamin WangSIGGRAPH 2022 · 被引用 48 次
- Extending Sparse Patterns to Improve Inverse Preconditioning on GPU ArchitecturesSergi Laut, Ricard Borrell, Marc CasasHPDC 2024 · 被引用 3 次
