SC2025Top-tier venue
Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures
Hansheng Wang, Dajun Huang, Gaoyuan Zou, Lu Shi, Xu Jiang, Xi Wu, Hancong Duan, Shaoshuai Zhang
Abstract
The 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Improving Tridiagonalization Performance on GPU ArchitecturesHansheng Wang, Zhekai Duan, Zitian Zhao, Siqi Wu et al.PPoPP 2025 · 6 citations
- Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor CoreShaoshuai Zhang, Ruchi Shah, Hiroyuki Ootomo, Rio Yokota et al.PPoPP 2023 · 5 citations
- W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUsJunmin Xiao, Yunfei Pang, Qing Xue, Chaoyang Shui et al.SC 2022 · 4 citations
- Faster proximal algorithms for matrix optimization using Jacobi-based eigenvalue methodsHamza Fawzi, Harry GoulbourneNeurIPS 2021 · 7 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
