CoopRT: Accelerating BVH Traversal for Ray Tracing via Cooperative Threads
Yavuz Selim Tozlu, Huiyang Zhou
摘要
Ray Tracing is a rendering technique to simulate the way light interacts with objects to create realistic images. It has become prominent thanks to the latest hardware support on Graphics Processing Units (GPUs), i.e., the Ray-Tracing (RT) unit, specially designed to accelerate ray tracing operations. Despite such hardware advances, ray tracing remains a performance bottleneck for high-performance graphics workloads, such as real-time path tracing (PT), which is an application of ray tracing where multiple bouncing rays are traced per pixel. The key reasons are (a) the costly Bounding Volume Hierarchy (BVH) traversal operation, and (b) low Single-Instruction-Multiple-Thread (SIMT) efficiency as the rays in the same warp deviate inevitably in their traversal paths.
In this work, we propose a novel architecture design for cooperative BVH traversal that exploits the parallelism present in the BVH traversal process. The key idea of our CoopRT scheme is to make use of the idle threads, either completely inactive when the ray tracing instruction is executed or partially idle due to early completion, to help the long running threads in the same warp. Specifically, we enable idle threads in a GPU warp to utilize their readily available traversal hardware to help traverse the BVH tree for the busy threads, therefore helping them finish their traversal much faster. This approach is implemented purely in hardware, requiring no changes to the programming model. We present our architecture design and show that it only involves small changes to the existing RT unit.
We evaluated CoopRT in Vulkan-sim, a cycle-level simulator, and observed up to 5.11x speedup over the baseline, with a geometric mean of 2.15x speedup at the cost of a moderate area overhead of 3.0% of the warp buffer in the RT unit. Using the energy-delay product, our CoopRT achieves an average of 2.29x improvement over the baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- RTNN: accelerating neighbor search using hardware ray tracingYuhao ZhuPPoPP 2022 · 被引用 43 次
- Vulkan-Sim: A GPU Architecture Simulator for Ray TracingMohammadreza Saed, Yuan-Hsi Chou, Lufei Liu, Tyler Nowicki 等MICRO 2022 · 被引用 30 次
- Intersection Prediction for Accelerated GPU Ray TracingLufei Liu, Wesley Chang, Francois Demoullin, Yuan-Hsi Chou 等MICRO 2021 · 被引用 24 次
- Generalizing Ray Tracing Accelerators for Tree Traversals on GPUsDongho Ha, Lufei Liu, Yuan-Hsi Chou, Seokjin Go 等MICRO 2024 · 被引用 16 次
相关 Paper
- TTP: A Hardware-Efficient Design for Precise Prefetching in Ray TracingYavuz Selim Tozlu, Anshul Naithani, Huiyang ZhouISCA 2026
- AQB8: Energy-Efficient Ray Tracing Accelerator through Multi-Level QuantizationYen-Chieh Huang, Chen-Pin Yang, Tsung Tai YehISCA 2025 · 被引用 1 次
- Treelet Prefetching For Ray TracingYuan-Hsi Chou, Tyler Nowicki, Tor M. AamodtMICRO 2023 · 被引用 12 次
- LibRTS: A Spatial Indexing Library by Ray TracingLiang Geng, Rubao Lee, Xiaodong ZhangPPoPP 2025 · 被引用 11 次
- Treelet Accelerated Ray Tracing on GPUsYuan-Hsi Chou, Tor M. AamodtASPLOS 2025 · 被引用 3 次
