LIBRA: Memory Bandwidth- and Locality-Aware Parallel Tile Rendering
Aurora Tomás, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio González
Abstract
The increasing demand for high-quality graphics requires a significant increase in computational power of modern GPUs. The common approach to follow is augmenting the number of compute units (i.e., shader cores). However, this can result in underutilized resources if the workload is not properly balanced. This is particularly challenging in Tile-Based Rendering (TBR) GPUs, the predominant architecture in mobile GPUs, running graphics applications due to limited per-tile workload. This work proposes parallel tile rendering to efficiently in-crease the computational capabilities of TBR GPUs. This solves the problem of not having enough work to utilize the additional compute units but causes memory-intensive applications to underperform due to the increased memory pressure. To this end, we introduce LIBRA, a parallel tile rendering architecture that includes a novel locality-aware approach to schedule tiles to Raster Units to evenly distribute memory requests during the rendering of each frame. This alleviates memory congestion, therefore, reducing memory access time. LIBRA leverages frame-to-frame coherence to predict the memory pressure of each tile of a frame without penalizing the hit ratio of the cache memories. Evaluations over a wide range of commercial gaming applications show that LIBRA reduces the average memory latency by 13.5% and achieves an average speedup of 20.9%. It also provides an 11.4% improvement in throughput (frames per second) and a total GPU energy reduction of 9.2%, while adding negligible overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38d3dc78-28dc-4d52-acba-0be32a64d8ebBuilds on4
- CHOPIN: Scalable Graphics Rendering in Multi-GPU Systems via Parallel Image CompositionXiaowei Ren, Mieszko LisHPCA 2021 · 14 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
- DTexL: Decoupled Raster Pipeline for Texture LocalityDiya Joseph, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2022 · 3 citations
Related papers
- LIBRA: A High-Accuracy, Cost-Aware, and Coordinated Multi-GPU Page PrefetcherXiangyue Huang, Yanan Guo, Yuanchao XuISCA 2026
- Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignBingyao Li, Jieming Yin, Youtao Zhang, Xulong TangMICRO 2021 · 33 citations
- VISTA: Optimizing GPU Scheduling through Versatile Locality-Aware Data SharingHajar Falahati, Negin Mahani, Adrián Cristal, Osman S. UnsalDAC 2025 · 1 citation
- Optimizing the Memory Hierarchy by Compositing Automatic Transformations on Computations and DataJie Zhao, Peng DiMICRO 2020 · 32 citations
- Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip ResourcesJagadish B. Kotra, Michael LeBeane, Mahmut T. Kandemir, Gabriel H. LohMICRO 2021 · 17 citations
