Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme
William S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim, Sri Hari Krishna Narayanan, Michel Schanen, Johannes Doerfert
摘要
Computing derivatives is key to many algorithms in scientific computing and machine learning such as optimization, uncertainty quantification, and stability analysis. Enzyme is a LLVM compiler plugin that performs reverse-mode automatic differentiation (AD) and thus generates high performance gradients of programs in languages including C/C++, Fortran, Julia, and Rust. Prior to this work, Enzyme and other AD tools were not capable of generating gradients of GPU kernels. Our paper presents a combination of novel techniques that make Enzyme the first fully automatic reversemode AD tool to generate gradients of GPU kernels. Since unlike other tools Enzyme performs automatic differentiation within a general-purpose compiler, we are able to introduce several novel GPU and AD-specific optimizations. To show the generality and efficiency of our approach, we compute gradients of five GPU-based HPC applications, executed on NVIDIA and AMD GPUs. All benchmarks run within an order of magnitude of the original program's execution time. Without GPU and AD-specific optimizations, gradients of GPU kernels either fail to run from a lack of resources or have infeasible overhead. Finally, we demonstrate that increasing the problem size by either increasing the number of threads or increasing the work per thread, does not substantially impact the overhead from differentiation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DR.JIT: a just-in-time compiler for differentiable renderingWenzel Jakob, Sébastien Speierer, Nicolas Roussel, Delio ViciniSIGGRAPH 2022 · 被引用 160 次
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel ConstructsWilliam S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo 等PPoPP 2023 · 被引用 27 次
- Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler AugmentationWilliam S. Moses, Sri Hari Krishna Narayanan, Ludger Paehler, Valentin Churavy 等SC 2022 · 被引用 25 次
- Aδ: autodiff for discontinuous programs - applied to shadersYuting Yang, Connelly Barnes, Andrew Adams, Adam FinkelsteinSIGGRAPH 2022 · 被引用 17 次
- AD for an Array Language with Nested ParallelismRobert Schenck, Ola Rønning, Troels Henriksen, Cosmin E. OanceaSC 2022 · 被引用 11 次
它引用的顶会 Paper3
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun 等ICLR 2020 · 被引用 479 次
- JAX MD: A Framework for Differentiable PhysicsSamuel S. Schoenholz, Ekin Dogus CubukNeurIPS 2020 · 被引用 195 次
- Instead of Rewriting Foreign Code for Machine Learning, Automatically Synthesize Fast GradientsWilliam S. Moses, Valentin ChuravyNeurIPS 2020 · 被引用 144 次
相关 Paper
- Descend: A Safe GPU Systems Programming LanguageBastian Köpcke, Sergei Gorlatch, Michel SteuwerPLDI 2024 · 被引用 7 次
- Locality-Aware Automatic Differentiation on the GPU for Mesh-Based ComputationsAhmed H. Mahmoud, Rahul Goel, Jonathan Ragan-Kelley, Justin SolomonSIGGRAPH 2026
- A simple differentiable programming languageMartín Abadi, Gordon D. PlotkinPOPL 2020 · 被引用 49 次
- Efficient Dual-Numbers Reverse AD via Well-Known Program TransformationsTom Smeding, Matthijs VákárPOPL 2023 · 被引用 10 次
- Provably correct, asymptotically efficient, higher-order reverse-mode automatic differentiationFaustyna Krawiec, Simon Peyton Jones, Neel Krishnaswami, Tom Ellis 等POPL 2022 · 被引用 27 次
