SC2025Top-tier venue
Caracal: A GPU-Resident Sparse LU Solver with Lightweight Fine-Grained Scheduling
Jie Ren, Tingxuan Zhong, Yuxi Hong, Guofeng Feng, Xincheng Wang, Weile Jia, Hatem Ltaief, David Elliot Keyes
Abstract
We address inefficiencies in task scheduling, memory management, and scalability in GPU-resident sparse LU factorization with a two-level approach of sequentially scheduled coarse-grained blocks containing multiple fine-grained blocks managed with a lightweight static scheduler enabling multi-stream parallelism. Additionally, we design an intelligent memory caching mechanism for the fine-grained scheduler, which retains frequently accessed data in GPU memory. To further enhance scalability, we introduce a distributed memory design that partitions the input matrix using a 1D block-cyclic distribution and optimizes inter-GPU communication via NVLink. The multi-GPU design reaches a computational throughput of 6.46 TFLOP/s on four A100 GPUs, demonstrating promising scalability. This is up to 7x speedup over the latest SuperLU_DIST with 3D communication, 94x speedup over PanguLU, 16x speedup over PasTiX, and 10x speedup over our own coarse-grained dynamic scheduling implementation while reaching up to 21% of the A100’s theoretical peak performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4c012a3b-d485-4b48-a5b4-2fe93896463eCited by top-tier papers1
Ask how each one uses itRelated papers
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin et al.DAC 2021 · 30 citations
- DAS-ILU: A Distributed Asynchronous Parallel ILU Factorization Based on Domain DecompositionFan Yuan, Shengguo Li, Xiaojian Yang, Yunqing Huang et al.SC 2025 · 1 citation
- End-to-End LU Factorization of Large Matrices on GPUsYang Xia, Peng Jiang, Gagan Agrawal, Rajiv RamnathPPoPP 2023 · 6 citations
- Spatula: A Hardware Accelerator for Sparse Matrix FactorizationAxel Feldmann, Daniel SánchezMICRO 2023 · 9 citations
- Symmetric Block-Cyclic Distribution: Fewer Communications Leads to Faster Dense Cholesky FactorizationOlivier Beaumont, Philippe Duchon, Lionel Eyraud-Dubois, Julien Langou et al.SC 2022 · 7 citations
