Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs
Rupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken, Maryam Mehri Dehnavi, Michael Garland, Michael Bauer
摘要
GPU architectures have continued to grow in complexity, with recent incarnations introducing increasingly powerful fixed-function units for matrix multiplication and data movement to accompany highly parallel general-purpose cores. To fully leverage these machines, software must use sophisticated schedules that maximally utilize all hardware resources. Since realizing such schedules is complex, both programmers and compilers routinely employ program transformations, such as software pipelining (SWP) and warp specialization (WS), to do so in practice. However, determining how best to use SWP and WS in combination is a challenging problem that is currently handled through a mix of brittle compilation heuristics and fallible human intuition, with little insight into the space of solutions. To remedy this situation, we introduce a novel formulation of SWP and WS as a joint optimization problem that can be solved holistically by off-the-shelf constraint solvers. We reify our approach in Twill, the first system that automatically derives optimal SWP and WS schedules for a large class of iterative programs. Twill is heuristic-free, easily extensible to new GPU architectures, and guaranteed to produce optimal schedules. We show that Twill can rediscover, and thereby prove optimal, the SWP and WS schedules manually developed by experts for Flash Attention on both the NVIDIA Hopper and Blackwell GPU architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia 等OSDI 2025 · 被引用 9 次
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 被引用 5 次
相关 Paper
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 被引用 13 次
- QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUsNicolás Meseguer, Daoxuan Xu, Yifan Sun, Michael Pellauer 等HPCA 2026
- A sparse iteration space transformation framework for sparse tensor algebraRyan Senanayake, Changwan Hong, Ziheng Wang, Amalee Wilson 等OOPSLA 2020 · 被引用 51 次
- Efficient automatic scheduling of imaging and vision pipelines for the GPULuke Anderson, Andrew Adams, Karima Ma, Tzu-Mao Li 等OOPSLA 2021 · 被引用 14 次
- Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand CompactionEunbi Jeong, Ipoom Jeong, Myung Kuk Yoon, Nam Sung KimHPCA 2025 · 被引用 2 次
