End-to-End LU Factorization of Large Matrices on GPUs
Yang Xia, Peng Jiang, Gagan Agrawal, Rajiv Ramnath
Abstract
LU factorization for sparse matrices is an important computing step for many engineering and scientific problems such as circuit simulation. There have been many efforts toward parallelizing and scaling this algorithm, which include the recent efforts targeting the GPUs. However, it is still challenging to deploy a complete sparse LU factorization workflow on a GPU due to high memory requirements and data dependencies. In this paper, we propose the first complete GPU solution for sparse LU factorization. To achieve this goal, we propose an out-of-core implementation of the symbolic execution phase, thus removing the bottleneck due to large intermediate data structures. Next, we propose a dynamic parallelism implementation of Kahn's algorithm for topological sort on the GPUs. Finally, for the numeric factorization phase, we increase the parallelism degree by removing the memory limits for large matrices as compared to the existing implementation approaches. Experimental results show that compared with an implementation modified from GLU 3.0, our out-of-core version achieves speedups of 1.13--32.65X. Further, our out-of-core implementation achieves a speedup of 1.2--2.2 over an optimized unified memory implementation on the GPU. Finally, we show that the optimizations we introduce for numeric factorization turn out to be effective.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ef348b03-d423-4e8e-bd9c-409e5d33f041Cited by top-tier papers2
- PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous SystemsXu Fu, Bingbin Zhang, Tengcheng Wang, Wenhao Li et al.SC 2023 · 23 citations
- Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU ClustersYida Li, Siwei Zhang, Yiduo Niu, Yang Du et al.PPoPP 2026 · 1 citation
Related papers
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin et al.DAC 2021 · 30 citations
- Accelerating Sparse LU Factorization with Density-Aware Adaptive Matrix Multiplication for Circuit SimulationTengcheng Wang, Wenhao Li, Haojie Pei, Yuying Sun et al.DAC 2023 · 23 citations
- Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct SolversAhmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov et al.SC 2022 · 4 citations
- Sparsified Preconditioned Conjugate Gradient Solver on GPUsDa Ma, Khalid Ahmad, Kazem Cheshmi, Hari Sundar et al.SC 2025 · 1 citation
- Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained SchedulingJie Ren, Tingxuan Zhong, Yuxi Hong, Guofeng Feng et al.SC 2025 · 1 citation
