Modular GPU Programming with Typed Perspectives
Manya Bansal, Daniel Sainati, Joseph W. Cutler, Saman P. Amarasinghe, Jonathan Ragan-Kelley
Abstract
To achieve peak performance on modern GPUs, one must balance two frames of mind: issuing instructions to individual threads to control their behavior, while simultaneously tracking the convergence of many threads acting in concert to perform collective operations like Tensor Core instructions. The tension between these two mindsets makes modular programming error prone. Functions that encapsulate collective operations, despite being called per-thread, must be executed cooperatively by groups of threads. In this work, we introduce Prism, a new GPU language that restores modularity while still giving programmers the low-level control over collective operations necessary for high performance. Our core idea is typed perspectives , which materialize, at the type level, the granularity at which the programmer is controlling the behavior of threads. We describe the design of Prism, implement a compiler for it, and lay its theoretical foundations in a core calculus called Bundl. We implement state-of-the-art GPU kernels in Prism and find that it offers programmers the safety guarantees needed to confidently write modular code without sacrificing performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94b2a00a-311b-407b-aa2d-93708ca70d21Builds on7
- Transfinite Iris: resolving an existential dilemma of step-indexed separation logicSimon Spies, Lennard Gäher, Daniel Gratzer, Joseph Tassarotti et al.PLDI 2021 · 32 citations
- Modeling and analyzing evaluation cost of CUDA kernelsStefan K. Muller, Jan HoffmannPOPL 2021 · 15 citations
- Descend: A Safe GPU Systems Programming LanguageBastian Köpcke, Sergei Gorlatch, Michel SteuwerPLDI 2024 · 7 citations
- Transfinite step-indexing for terminationSimon Spies, Neel Krishnaswami, Derek DreyerPOPL 2021 · 7 citations
- Asynchronous effectsDanel Ahman, Matija PretnarPOPL 2021 · 6 citations
Related papers
- Graphene: An IR for Optimized Tensor Computations on GPUsBastian Hagedorn, Bin Fan, Hanfeng Chen, Cris Cecka et al.ASPLOS 2023 · 30 citations
- SIMT-Step Execution: A Flexible Operational Semantics for GPU Subgroup BehaviorZheyuan Chen, Naomi Rehman, Guido Martínez, Tyler SorensenPLDI 2026 · 2 citations
- Kuiper: Correct and Efficient GPU Programming with Dependent Types and Separation LogicGuido Martínez, Bastian Köpcke, Jonás Fiala, Gabriel Ebner et al.PLDI 2026 · 2 citations
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- TensorPrism: Rethinking Sparse High-Order Tensor Acceleration via Co-Occurrence GraphFangzhou Ye, Shilin Tian, Amir Ghazizadeh Ahsaei, Hao ZhengISCA 2026
