RTSpMSpM: Harnessing Ray Tracing for Efficient Sparse Matrix Computations
Hongrui Zhang, Yunan Zhang, Hung-Wei Tseng
摘要
The significance of sparse matrix algebra pushes the development of sparse matrix accelerators.Despite the general reception of using hardware accelerators to address application demands and the convincement of substantial performance gain, integrating heterogeneous hardware accelerators always introduces manufacturing costs on new hardware and challenges the system interconnects.Inspired by the similar algorithmic behaviors between ray tracing and sparse matrix problems, this paper exploits ray tracing hardware.This increasingly popular accelerator has become the standard in modern GPU architectures to address the demand for virtual/augmented/mixed reality applications.We propose an algorithm that maps the most classical sparse matrix multiplication problem (SpMSpM) as a ray tracing problem.By implementing the proposed algorithm using a modern ray tracing programming framework, the resulting program, SW-RTSpMSpM can leverage the accelerated intersection test unit in commercialized GPUs and reveals 1.85× speedup over the state-of-the-art SpMSpM library.Despite the "similarities" in both problems enabling the potential of accelerating SpMSpM using existing ray tracing hardware, the "difference" between these two algorithms leaves room for performance gain with architectural optimizations.The insights from evaluating SW-RTSpMSpM guide the proposal of RT+SpMSpM, where we present two major architectural optimizations that (1) allow simultaneous computation on multiplications along with intersection tests and (2) ray mapping and scheduling and a small row accumulation engine to leverage Gustavson's dataflow and eliminate the reliance of additional shader functions and redundant memory operations.The resulting RT+SpMSpM only introduces 0.2% area overhead to a modern high-end ray-tracing-hardware-equipped GPU and improves SW-RTSpMSpM by 1.66×, achieving 3.06× speedup over software library and 80% performance per area compared to a stateof-the-art SpMSpM accelerator, without losing the capability in supporting ray tracing.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- GRTX: Efficient Ray Tracing for 3D Gaussian-Based RenderingJunseo Lee, Sangyun Jeon, Jungi Lee, Junyong Park 等HPCA 2026 · 被引用 2 次
- TTP: A Hardware-Efficient Design for Precise Prefetching in Ray TracingYavuz Selim Tozlu, Anshul Naithani, Huiyang ZhouISCA 2026
相关 Paper
- Intersection Prediction for Accelerated GPU Ray TracingLufei Liu, Wesley Chang, Francois Demoullin, Yuan-Hsi Chou 等MICRO 2021 · 被引用 24 次
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- LibRTS: A Spatial Indexing Library by Ray TracingLiang Geng, Rubao Lee, Xiaodong ZhangPPoPP 2025 · 被引用 11 次
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta 等SC 2024 · 被引用 15 次
- pSyncPIM: Partially Synchronous Execution of Sparse Matrix Operations for All-Bank PIM ArchitecturesDaehyeon Baek, Soojin Hwang, Jaehyuk HuhISCA 2024 · 被引用 18 次
