WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp Specialization
Neal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. Keckler
摘要
Graphics processing units (GPUs) are an important class of parallel processors that offer high compute throughput and memory bandwidth. GPUs are used in a variety of important computing domains, such as machine learning, high performance computing, sparse linear algebra, autonomous vehicles, and robotics. However, some applications from these domains can underperform due to sensitivity to memory latency and bandwidth. Some of this sensitivity can be reduced by better overlapping memory access with compute. Current GPUs often leverage pipeline parallelism in the form of warp specialization to enable better overlap. However, current warp specialization support on GPUs is limited in three ways. First, warp specialization is a complex and manual program transformation that is out of reach for many applications and developers. Second, it is limited to coarse-grained transfers between global memory and the shared memory scratchpad (SMEM); fine-grained memory access patterns are not well supported. Finally, the GPU hardware is unaware of the pipeline parallelism expressed by the programmer, and is unable to take advantage of this information to make better decisions at runtime. In this paper we introduce WASP, hardware and compiler support for warp specialization that addresses these limitations. WASP enables fine-grained streaming and gather memory access patterns through the use of warp-level register file queues and hardware-accelerated address generation. Explicit warp to pipeline stage naming enables the GPU to be aware of pipeline parallelism, which WASP capitalizes on by designing pipeline-aware warp mapping, register allocation, and scheduling. Finally, we design and implement a compiler that can automatically generate warp specialized kernels, reducing programmer burden. Overall, we find that runtime performance can be improved on a variety of important applications by an average of 47% over a modern GPU baseline.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia 等OSDI 2025 · 被引用 9 次
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones 等SC 2025 · 被引用 3 次
- LIBRA: Memory Bandwidth- and Locality-Aware Parallel Tile RenderingAurora Tomás, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2024 · 被引用 1 次
- KPerfIR: Towards a Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI WorkloadsYue Guan, Yuanwei Fang, Keren Zhou, Corbin Robeck 等OSDI 2025
相关 Paper
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken 等OSDI 2026 · 被引用 9 次
- The Sparsity-Aware LazyGPU ArchitectureChangxi Liu, Miao Yu, Yifan Sun, Trevor E. CarlsonISCA 2025 · 被引用 1 次
- SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsShinnung Jeong, Liam Paul Cooper, Ju Min Lee, Heelim Choi 等HPCA 2025 · 被引用 2 次
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 被引用 2 次
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 被引用 7 次
