VISTA: Optimizing GPU Scheduling through Versatile Locality-Aware Data Sharing
Hajar Falahati, Negin Mahani, Adrián Cristal, Osman S. Unsal
Abstract
Graphics Processing Units (GPUs) play a pivotal role in high-performance computing by utilizing the massive parallelism of concurrent thread execution to enhance processing efficiency, known as thread-level parallelism (TLP). However, effectively managing the significant memory access generated by a large number of concurrent threads presents a significant challenge, affecting both performance and energy consumption. However, previous GPU scheduling methods often overlook the significant potential of data sharing among non-adjacent warps or Cooperative Thread Arrays (CTAs), thereby limiting their effectiveness in reducing unnecessary memory accesses through supporting data sharing. Our observation underscores the untapped potential of non-adjacent data sharing, motivating our work to optimize schedules more effectively. In this paper, we present VISTA, a smart locality-aware GPU scheduler that dynamically identifies data locality patterns at both the CTA and warp levels, leveraging this runtime information to make intelligent scheduling decisions and improving memory efficiency and overall performance. VISTA significantly reduces unnecessary memory accesses. Simulation results show a 48.1% improvement in performance and a 51.8% reduction in energy consumption for memory-intensive GPGPU applications, compared to the baseline, with negligible hardware overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- DTexL: Decoupled Raster Pipeline for Texture LocalityDiya Joseph, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2022 · 3 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU CommunicationsJiajian Zhang, Fangyu Wu, Hai Jiang, Qiufeng Wang et al.USENIX ATC 2025 · 1 citation
- R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUsDongho Ha, Yunho Oh, Won Woo RoISCA 2023 · 9 citations
- Snake: A Variable-length Chain-based Prefetching for GPUsSaba Mostofi, Hajar Falahati, Negin Mahani, Pejman Lotfi-Kamran et al.MICRO 2023 · 11 citations
