SC2022Top-tier venue
AD for an Array Language with Nested Parallelism
Robert Schenck, Ola Rønning, Troels Henriksen, Cosmin E. Oancea
Abstract
We present a technique for applying (forward and) reversemode automatic differentiation (AD) on a non-recursive secondorder functional array language that supports nested parallelism and is primarily aimed at efficient GPU execution.
The key idea is to eliminate the need for a "tape" by relying on redundant execution to bring into each new scope all program variables that may be needed by the differentiated code. Efficient execution is enabled by the observation that perfectly-nested scopes do not introduce re-execution, and such perfect nests are produced by known compiler transformations, e.g., flattening. Our technique differentiates loops and bulk-parallel operators-such as map, reduce, histogram, scan, scatter-by specific rewrite rules, and aggressively optimizes the resulting nested-parallel code. We report an experimental evaluation that compares with established AD solutions and demonstrates competitive performance on nine common benchmarks from recent applied AD literature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 202b93a6-7d6e-448a-82d0-3040837e52d3Cited by top-tier papers4
- Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler AugmentationWilliam S. Moses, Sri Hari Krishna Narayanan, Ludger Paehler, Valentin Churavy et al.SC 2022 · 25 citations
- Efficient Dual-Numbers Reverse AD via Well-Known Program TransformationsTom Smeding, Matthijs VákárPOPL 2023 · 10 citations
- Efficient CHADTom Smeding, Matthijs VákárPOPL 2024 · 6 citations
- Verifying Array Properties in Pure Data-Parallel ProgramsNikolaj Hey Hinnerskov, Robert Schenck, Cosmin E. OanceaPLDI 2026 · 1 citation
Builds on3
- Instead of Rewriting Foreign Code for Machine Learning, Automatically Synthesize Fast GradientsWilliam S. Moses, Valentin ChuravyNeurIPS 2020 · 144 citations
- Reverse-mode automatic differentiation and optimization of GPU kernels via enzymeWilliam S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim et al.SC 2021 · 50 citations
- Compiling generalized histograms for GPUTroels Henriksen, Sune Hellfritzsch, Ponnuswamy Sadayappan, Cosmin E. OanceaSC 2020 · 10 citations
Related papers
- ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct IndexingShuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao et al.PPoPP 2026
- Provably correct, asymptotically efficient, higher-order reverse-mode automatic differentiationFaustyna Krawiec, Simon Peyton Jones, Neel Krishnaswami, Tom Ellis et al.POPL 2022 · 27 citations
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 1 citation
- A simple differentiable programming languageMartín Abadi, Gordon D. PlotkinPOPL 2020 · 49 citations
- Locality-Aware Automatic Differentiation on the GPU for Mesh-Based ComputationsAhmed H. Mahmoud, Rahul Goel, Jonathan Ragan-Kelley, Justin SolomonSIGGRAPH 2026
