A Diagonal Block Memory-Aware Polynomial Preconditioner for Linear and Eigenvalue Solvers
Xiaojian Yang, Yuhui Ni, Fan Yuan, Shengguo Li, Dezun Dong, Chuanfu Xu, Haipeng Jia, Jie Liu
Abstract
Krylov subspace methods are widely used in scientific computing to solve large sparse linear systems and eigenvalue problems. Their performance bottleneck is often dominated by high-order matrix-power kernels (MPK), especially in polynomial preconditioners that must scale to millions or billions of variables. We present Diagonal Block MPK (DBMPK), a lightweight and parallel-friendly optimization that partitions the input matrix into diagonal blocks and off-diagonal regions. This design enables efficient intra-block data reuse and eliminates inter-block dependencies. It improves cache locality, parallelism, and reduces preprocessing overheads, compared to existing techniques. Our evaluation on x86 and Arm HPC platforms shows that DBMPK improves MPK performance by 26.6%-38.4%. When applied to polynomial preconditioners for linear systems and eigenvalue problems, it achieves consistent end-to-end speedups of 18.6%-34.0%, including in weak scaling tests on 128 nodes, demonstrating strong scalability and practical impact.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 25adabe3-eb7b-4162-a347-d2c36ac01b69Related papers
- Me-MPK: Accelerating Krylov Subspace Solvers via Memory-efficient Matrix-Power KernelHaozhong Qiu, Chuanfu Xu, Jianbin Fang, Shengguo Li et al.DAC 2025 · 1 citation
- Sparsified Preconditioned Conjugate Gradient Solver on GPUsDa Ma, Khalid Ahmad, Kazem Cheshmi, Hari Sundar et al.SC 2025 · 1 citation
- A submatrix-based method for approximate matrix function evaluation in the quantum chemistry code CP2KMichael Lass, Robert Schade, Thomas D. Kühne, Christian PlesslSC 2020 · 7 citations
- Solving Linear Systems on a GPU with Hierarchically Off-Diagonal Low-Rank ApproximationsChao Chen, Per-Gunnar MartinssonSC 2022 · 4 citations
- Exploiting Hierarchical Parallelism and Reusability in Tensor Kernel Processing on Heterogeneous HPC SystemsYuedan Chen, Guoqing Xiao, M. Tamer Özsu, Zhuo Tang et al.ICDE 2022 · 7 citations
