Task-Based Tensor Computations on Modern GPUs
Rohan Yadav, Michael Garland, Alex Aiken, Michael Bauer
Abstract
Domain-specific, fixed-function units are becoming increasingly common in modern processors. As the computational demands of applications evolve, the capabilities and programming interfaces of these fixed-function units continue to change. NVIDIA’s Hopper GPU architecture contains multiple fixed-function units per compute unit, including an asynchronous data movement unit (TMA) and an asynchronous matrix multiplication unit (Tensor Core). Efficiently utilizing these units requires a fundamentally different programming style than previous architectures; programmers must now develop warp-specialized kernels that orchestrate producer consumer pipelines between the asynchronous units. To manage the complexity of programming these new architectures, we introduce Cypress, a task-based programming model with sequential semantics. Cypress programs are a set of designated functions called tasks that operate on tensors and are free of communication and synchronization. Cypress programs are bound to the target machine through a mapping specification that describes where tasks should run and in which memories tensors should be materialized. We present a compiler architecture that lowers Cypress programs into CUDA programs that perform competitively with expert-written codes. Cypress achieves 0.88x-1.06x the performance of cuBLAS on GEMM, and between 0.80x-0.98x the performance of the currently best-known Flash Attention implementation while eliminating all aspects of explicit data movement and asynchronous computation from application code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 819fc479-0a03-4333-a703-12a56ec33e67Cited by top-tier papers4
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken et al.OSDI 2026 · 9 citations
- Modular GPU Programming with Typed PerspectivesManya Bansal, Daniel Sainati, Joseph W. Cutler, Saman P. Amarasinghe et al.PLDI 2026
- Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationJaehyeong Park, Youngchan Kim, Haechan An, Gieun Jeong et al.ASPLOS 2026
- TileLang: Bridge Programmability and Performance in Modern Neural KernelsLei Wang, Yu Cheng, Yining Shi, Zhiwen Mo et al.ICLR 2026
Builds on6
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- egg: Fast and extensible equality saturationMax Willsey, Chandrakana Nandi, Yisu Remy Wang, Oliver Flatt et al.POPL 2021 · 170 citations
- Exocompilation for productive programming of hardware acceleratorsYuka Ikarashi, Gilbert Louis Bernstein, Alex Reinking, Hasan Genc et al.PLDI 2022 · 56 citations
- Graphene: An IR for Optimized Tensor Computations on GPUsBastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka et al.ASPLOS 2023 · 30 citations
Related papers
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia et al.OSDI 2025 · 9 citations
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 8 citations
- QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUsNicolás Meseguer, Daoxuan Xu, Yifan Sun, Michael Pellauer et al.HPCA 2026
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- Descend: A Safe GPU Systems Programming LanguageBastian Köpcke, Sergei Gorlatch, Michel SteuwerPLDI 2024 · 7 citations
