Modular GPU Programming with Typed Perspectives
Manya Bansal, Daniel Sainati, Joseph W. Cutler, Saman P. Amarasinghe, Jonathan Ragan-Kelley
摘要
To achieve peak performance on modern GPUs, one must balance two frames of mind: issuing instructions to individual threads to control their behavior, while simultaneously tracking the convergence of many threads acting in concert to perform collective operations like Tensor Core instructions. The tension between these two mindsets makes modular programming error prone. Functions that encapsulate collective operations, despite being called per-thread, must be executed cooperatively by groups of threads. In this work, we introduce Prism, a new GPU language that restores modularity while still giving programmers the low-level control over collective operations necessary for high performance. Our core idea is typed perspectives , which materialize, at the type level, the granularity at which the programmer is controlling the behavior of threads. We describe the design of Prism, implement a compiler for it, and lay its theoretical foundations in a core calculus called Bundl. We implement state-of-the-art GPU kernels in Prism and find that it offers programmers the safety guarantees needed to confidently write modular code without sacrificing performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Transfinite Iris: resolving an existential dilemma of step-indexed separation logicSimon Spies, Lennard Gäher, Daniel Gratzer, Joseph Tassarotti 等PLDI 2021 · 被引用 32 次
- Modeling and analyzing evaluation cost of CUDA kernelsStefan K. Muller, Jan HoffmannPOPL 2021 · 被引用 15 次
- Descend: A Safe GPU Systems Programming LanguageBastian Köpcke, Sergei Gorlatch, Michel SteuwerPLDI 2024 · 被引用 7 次
- Transfinite step-indexing for terminationSimon Spies, Neel Krishnaswami, Derek DreyerPOPL 2021 · 被引用 7 次
- Asynchronous effectsDanel Ahman, Matija PretnarPOPL 2021 · 被引用 6 次
相关 Paper
- Graphene: An IR for Optimized Tensor Computations on GPUsBastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka 等ASPLOS 2023 · 被引用 30 次
- SIMT-Step Execution: A Flexible Operational Semantics for GPU Subgroup BehaviorZheyuan Chen, Naomi Rehman, Guido Martínez, Tyler SorensenPLDI 2026 · 被引用 2 次
- Kuiper: Correct and Efficient GPU Programming with Dependent Types and Separation LogicGuido Martínez, Bastian Köpcke, Jonás Fiala, Gabriel Ebner 等PLDI 2026 · 被引用 2 次
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 被引用 5 次
- TensorPrism: Rethinking Sparse High-Order Tensor Acceleration via Co-Occurrence GraphFangzhou Ye, Shilin Tian, Amir Ghazizadeh Ahsaei, Hao ZhengISCA 2026
