SC2023Top-tier venue
Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU Clusters
Yang Liu, Nan Ding, Piyush Sao, Samuel Williams, Xiaoye Sherry Li
Abstract
This paper presents a unified communication optimization framework for sparse triangular solve (SpTRSV) algorithms on CPU and GPU clusters. The framework builds upon a 3D communicationavoiding (CA) layout of 𝑃 𝑥 × 𝑃 𝑦 × 𝑃 𝑧 processes that divides a sparse matrix into 𝑃 𝑧 submatrices, each handled by a 𝑃 𝑥 × 𝑃 𝑦 2D grid with block-cyclic distribution. We propose three communication optimization strategies: First, a new 3D SpTRSV algorithm is developed, which trades the inter-grid communication and synchronization with replicated computation. This design requires only one inter-grid synchronization, and the inter-grid communication is efficiently implemented with sparse allreduce operations. Second, broadcast and reduction communication trees are used to reduce message latency of the intra-grid 2D communication on CPU clusters. Finally, we leverage GPU-initiated one-sided communication to implement the communication trees on GPU clusters. With these nested inter-and intra-grid communication optimization strategies, the proposed 3D SpTRSV algorithm can attain up to 3.45x speedups compared to the baseline 3D SpTRSV algorithm using up to 2048 Cori Haswell CPU cores. In addition, the proposed GPU 3D Sp-TRSV algorithm can achieve up to 6.5x speedups compared to the proposed CPU 3D SpTRSV algorithm with 𝑃 𝑧 up to 64. Finally it is remarkable that the proposed GPU 3D SpTRSV can scale to 256 GPUs using the Perlmutter system while the existing 2D SpTRSV algorithm can only scale up to 4 GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68270095-0158-4e18-b3a2-36de7ed23104Cited by top-tier papers3
- PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous SystemsXu Fu, Bingbin Zhang, Tengcheng Wang, Wenhao Li et al.SC 2023 · 23 citations
- A Workflow Roofline Model for End-to-End Workflow Performance AnalysisNan Ding, Brian Austin, Yang Liu, Neil Mehta et al.SC 2024 · 6 citations
- Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU ClustersYida Li, Siwei Zhang, Yiduo Niu, Yang Du et al.PPoPP 2026 · 1 citation
Related papers
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- Trivance: Latency-Optimal AllReduce by Shortcutting Multiport NetworksAnton Juerss, Vamsi Addanki, Stefan SchmidSIGCOMM 2026 · 1 citation
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU ClustersJames D. Trotter, Sinan Ekmekçibasi, Dogan Sagbili, Johannes Langguth et al.SC 2025 · 2 citations
- Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix MultiplicationIsuru Ranawaka, Md Taufique Hussain, Charles Block, Gerasimos Gerogiannis et al.SC 2024 · 5 citations
