Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous Architectures
Hansheng Wang, Dajun Huang, Gaoyuan Zou, Lu Shi, Xu Jiang, Xi Wu, Hancong Duan, Shaoshuai Zhang
摘要
The 2-stage eigenvalue decomposition (EVD) method outperforms conventional 1-stage method on GPUs and heterogeneous architectures, especially when eigenvectors are not required. However, its performance advantage diminishes when performing back transformation to obtain eigenvectors. To address this, we propose two key solutions: 1) replacing BLAS3 operations with BLAS2 operations during the bulge-chasing back transformation for better performance, and 2) reordering the back transformation workflow from a backward pattern to a new parallelism-driven pattern to hide divide-and-conquer latency, at the cost of one additional GEMM computation. Experimentally, the proposed back transformation algorithm demonstrates significant performance improvements, outperforming the SOTA implementation in MAGMA by an average factor of 3.58x. For complete FP64 precision symmetric EVD with eigenvectors, the proposed algorithm, incorporating both solutions, surpasses the SOTA implementations in MAGMA and cuSOLVER by average factors of 2.62x and 2.21x, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Improving Tridiagonalization Performance on GPU ArchitecturesHansheng Wang, Zhekai Duan, Zitian Zhao, Siqi Wu 等PPoPP 2025 · 被引用 6 次
- Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor CoreShaoshuai Zhang, Ruchi Shah, Hiroyuki Ootomo, Rio Yokota 等PPoPP 2023 · 被引用 5 次
- W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUsJunmin Xiao, Yunfei Pang, Qing Xue, Chaoyang Shui 等SC 2022 · 被引用 4 次
- Faster proximal algorithms for matrix optimization using Jacobi-based eigenvalue methodsHamza Fawzi, Harry GoulbourneNeurIPS 2021 · 被引用 7 次
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian 等SC 2025 · 被引用 4 次
