TTP: A Hardware-Efficient Design for Precise Prefetching in Ray Tracing
Yavuz Selim Tozlu, Anshul Naithani, Huiyang Zhou
摘要
Ray tracing (RT) is a 3D graphics technique that offers highly realistic visuals. It is becoming prominent and accessible as GPU vendors have integrated dedicated ray tracing acceleration hardware. However, tracing millions of rays through 3D scenes consisting of high numbers of triangles in real time is challenging and requires expensive hardware. The main bottleneck in RT workloads is the expensive Bounding Volume Hierarchy (BVH) traversal task, which is a large tree structure that encodes the 3D scene. BVH traversal is a memory-bound problem, as the GPU threads spend most of their time reading tree node data from memory.
In this work, we attack the memory latency bottleneck of ray tracing through prefetching. We propose a novel hardware prefetcher, named Tree Traversal Prefetcher (TTP), for ray tracing. The main idea is to leverage the existing tree traversal stack in the RT units for highly accurate prefetching. In particular, TTP prefetches nodes using the addresses already available on the hardware traversal stacks of each thread. For DFS (Depth-first search) based traversal, prefetches are generated when nodes are being popped consecutively from the traversal stack, potentially corresponding to upward traversal through the tree.
We evaluate TTP on a cycle-level simulator, Vulkan-sim 2.0, and show that it achieves 1.48x speedup on average (up to 1.89x) compared to the baseline, with nearly negligible hardware overhead. TTP achieves 98.92% average L1 accuracy, which is the ratio of the prefetched blocks being actually referenced by demand loads. The coverage, computed as the ratio of L1 miss reduction over baseline L1 misses, is 31.54%, correlating well with the achieved speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- Vulkan-Sim: A GPU Architecture Simulator for Ray TracingMohammadreza Saed, Yuan-Hsi Chou, Lufei Liu, Tyler Nowicki 等MICRO 2022 · 被引用 30 次
- Intersection Prediction for Accelerated GPU Ray TracingLufei Liu, Wesley Chang, Francois Demoullin, Yuan-Hsi Chou 等MICRO 2021 · 被引用 24 次
- Generalizing Ray Tracing Accelerators for Tree Traversals on GPUsDongho Ha, Lufei Liu, Yuan-Hsi Chou, Seokjin Go 等MICRO 2024 · 被引用 16 次
- Treelet Prefetching For Ray TracingYuan-Hsi Chou, Tyler Nowicki, Tor M. AamodtMICRO 2023 · 被引用 12 次
相关 Paper
- CoopRT: Accelerating BVH Traversal for Ray Tracing via Cooperative ThreadsYavuz Selim Tozlu, Huiyang ZhouISCA 2025 · 被引用 2 次
- AQB8: Energy-Efficient Ray Tracing Accelerator through Multi-Level QuantizationYen-Chieh Huang, Chen-Pin Yang, Tsung Tai YehISCA 2025 · 被引用 1 次
- Extending GPU Ray-Tracing Units for Hierarchical Search AccelerationAaron Barnes, Fangjia Shen, Timothy G. RogersMICRO 2024 · 被引用 10 次
- Treelet Accelerated Ray Tracing on GPUsYuan-Hsi Chou, Tor M. AamodtASPLOS 2025 · 被引用 3 次
- LibRTS: A Spatial Indexing Library by Ray TracingLiang Geng, Rubao Lee, Xiaodong ZhangPPoPP 2025 · 被引用 11 次
