Trojan Horse: Aggregate-and-Batch for Scaling Up Sparse Direct Solvers on GPU Clusters
Yida Li, Siwei Zhang, Yiduo Niu, Yang Du, Qingxiao Sun, Zhou Jin, Weifeng Liu
Abstract
Sparse direct solvers are critical building blocks in a range of scientific applications on heterogeneous supercomputers. However, existing sparse direct solvers have not been able to well leverage the high bandwidth and floating-point performance of modern GPUs. The primary challenges are twofold:
(1) the absence of a mechanism for aggregating small tasks to saturate the GPU, and (2) the lack of a mechanism for executing a diverse set of small tasks in batch mode on a single GPU.
We in this paper propose a strategy called Trojan Horse, which significantly enhances the execution efficiency of sparse direct solvers on GPU clusters. This mechanism divides each process's work into two stages: Aggregate (with two modules Prioritizer and Container) and Batch (with two modules Collector and Executor). In the Aggregate stage, a process first assesses the urgency of the input tasks through the Prioritizer module, and based on their priority, sends them to the Collector module or the Container module. In the batch stage, the Collector module receives high-priority heterogeneous tasks from the Prioritizer module and retrieves
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a80d1cd5-264e-4e82-a303-db74638271edBuilds on15
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsYuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song et al.PPoPP 2022 · 66 citations
- Task bench: a parameterized benchmark for evaluating parallel runtime performanceElliott Slaughter, Wei Wu, Yuankun Fu, Legend Brandenburg et al.SC 2020 · 51 citations
- MTM: Rethinking Memory Profiling and Migration for Multi-Tiered Large MemoryJie Ren, Dong Xu, Junhee Ryu, Kwangsik Shin et al.EuroSys 2024 · 31 citations
- SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUsJianqi Zhao, Yao Wen, Yuchen Luo, Zhou Jin et al.DAC 2021 · 30 citations
- PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous SystemsXu Fu, Bingbin Zhang, Tengcheng Wang, Wenhao Li et al.SC 2023 · 23 citations
Related papers
- Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct SolversAhmad Abdelfattah, Pieter Ghysels, Wajih Boukaram, Stanimire Tomov et al.SC 2022 · 4 citations
- A fast work-efficient SSSP algorithm for GPUsKai Wang, Don Fussell, Calvin LinPPoPP 2021 · 21 citations
- Extending Sparse Patterns to Improve Inverse Preconditioning on GPU ArchitecturesSergi Laut, Ricard Borrell, Marc CasasHPDC 2024 · 3 citations
- Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU ClustersYang Liu, Nan Ding, Piyush Sao, Samuel Williams et al.SC 2023 · 8 citations
- Orchestrating Large-Scale SpGEMMs using Dynamic Block Distribution and Data Transfer Minimization on Heterogeneous SystemsTaehyeong Park, Seokwon Kang, Myung-Hwan Jang, Sang-Wook Kim et al.ICDE 2023 · 6 citations
