SC2022Top-tier venue
W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUs
Junmin Xiao, Yunfei Pang, Qing Xue, Chaoyang Shui, Ke Meng, Hui Ma, Mingyi Li, Xiaoyang Zhang, Guangming Tan
Abstract
As a basic matrix factorization operation, Singular Value Decomposition (SVD) is widely used in diverse domains. In real-world applications, the computational bottleneck of matrix factorization is on small matrices, and many GPU-accelerated batched SVD algorithms have been developed recently for higher performance. However, these algorithms failed to achieve both high data locality and convergence speed, because they are size-sensitive. In this work, we propose a novel W-cycle SVD to accelerate the batched one-sided Jacobi SVD on GPUs. The W-cycle SVD, which is size-oblivious, successfully exploits the data reuse and ensures the optimal convergence speed for batched SVD. Further, we present the efficient batched kernel design, and propose a tailoring strategy based on auto-tuning to improve the batched matrix multiplication in SVDs. The evaluation demonstrates that the proposed algorithm achieves 2.6∼10.2× speedup over the state-of-the-art cuSOLVER. In a real-world data assimilation application, our algorithm achieves 2.73∼3.09× speedup compared with MAGMA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3ab403b5-e3e5-4419-bd09-19dda7fd0df8Related papers
- Faster proximal algorithms for matrix optimization using Jacobi-based eigenvalue methodsHamza Fawzi, Harry GoulbourneNeurIPS 2021 · 7 citations
- Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUsJunmin Xiao, Chaoyang Shui, Di Cai, Kangyu Wang et al.SC 2023 · 1 citation
- Rethinking Back Transformation in 2-stage Eigenvalue Decomposition on Heterogeneous ArchitecturesHansheng Wang, Dajun Huang, Gaoyuan Zou, Lu Shi et al.SC 2025 · 2 citations
- HeteroSVD: Efficient SVD Accelerator on Versal ACAP with Algorithm-Hardware Co-DesignXinya Luan, Zhe Lin, Kai Shi, Jianwang Zhai et al.DAC 2025 · 1 citation
- Towards Singular Value Decomposition for Rank-Deficient Matrices: An Efficient and Accurate Algorithm on GPU ArchitecturesLu Shi, Weiwei Xu, Shaoshuai ZhangPPoPP 2026 · 1 citation
