DTexL: Decoupled Raster Pipeline for Texture Locality
Diya Joseph, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio González
Abstract
Contemporary GPU architectures have multiple shader cores and a scheduler that distributes work (threads) among them, focusing on load balancing. These load balancing techniques favor thread distributions that are detrimental to texture memory locality for graphics applications in the L1 Texture Caches. Texture memory accesses make up the majority of the traffic to the memory hierarchy in typical low power graphics architectures. This paper focuses on improving the L1 Texture cache locality by focusing on a new workload scheduler by exploring various methods to group the threads, assign the groups to shader cores and also to reorder threads without violating the correctness of the pipeline. To overcome the resulting load imbalance, we also propose a minor modification in the GPU architecture that helps translate the improvement in cache locality to an improvement in the GPU’s performance. We propose DTexL that envelops these ideas and evaluate it over a benchmark suite of ten commercial games, to obtain a 46.8% decrease in L2 Accesses, a 19.3% increase in performance and a 6.3% decrease in total GPU energy. All this with a negligible overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Analyzing and Leveraging Decoupled L1 Caches in GPUsMohamed Assem Ibrahim, Onur Kayiran, Yasuko Eckert, Gabriel H. Loh et al.HPCA 2021 · 30 citations
- VISTA: Optimizing GPU Scheduling through Versatile Locality-Aware Data SharingHajar Falahati, Negin Mahani, Adrián Cristal, Osman S. UnsalDAC 2025 · 1 citation
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- Mitigating GPU Core Partitioning Performance EffectsAaron Barnes, Fangjia Shen, Timothy G. RogersHPCA 2023 · 18 citations
- Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsBingyao Li, Yueqi Wang, Xulong TangDAC 2023 · 9 citations
