Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations
Ahmed H. Mahmoud, Rahul Goel, Jonathan Ragan-Kelley, Justin Solomon
Abstract
Fig. 1. We introduce a GPU system for efficient automatic differentiation of computations defined on triangle meshes that exploits locality and sparsity in mesh-based workloads. Using our system, users specify only the energy terms of their application while our system computes gradients, sparse Hessians, and Jacobians automatically and efficiently on the GPU. Here, we use a Newton solver for large-scale elastic shell simulation to simulate ≈ 700 Spot cows (≈2.1M vertices in total) falling to the ground where derivative computation accounts for only 12.2% of the total runtime.
We present a GPU-based system for automatic differentiation (AD) of functions defined on triangle meshes, designed to exploit the locality and sparsity in mesh-based computation. Our system evaluates derivatives using perelement forward-mode AD, confining all computation to registers and shared memory and assembling global gradients, sparse Jacobians, and sparse Hessians directly on the GPU. By avoiding global computation graphs, intermediate buffers, and device-host synchronization, our approach minimizes memory traffic and enables efficient differentiation under both static and dynamically changing sparsity. Our programming model lets users express energy terms over mesh neighborhoods, while our system automatically manages parallel execution, derivative propagation, sparse assembly, and matrix-free operations such as Hessian-vector products. Our system supports both scalar-and vector-valued objectives, dynamic interaction-driven sparsity updates, and seamless integration with external GPU sparse linear solvers. We evaluate our system on applications including elastic and cloth simulation, surface parameterization, mesh smoothing, frame field design, ARAP deformation, and spherical manifold optimization. Across these tasks, our system consistently outperforms state-of-the-art differentiation frameworks, including PyTorch, JAX, Warp, Dr.JIT, EnzymeAD, and Thallo. We
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0bc5465-17cd-4100-98a4-5c906e4dfc83Builds on8
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Incremental potential contact: intersection-and inversion-free, large-deformation dynamicsMinchen Li, Zachary Ferguson, Teseo Schneider, Timothy R. Langlois et al.SIGGRAPH 2020 · 320 citations
- DR.JIT: a just-in-time compiler for differentiable renderingWenzel Jakob, Sébastien Speierer, Nicolas Roussel, Delio ViciniSIGGRAPH 2022 · 160 citations
- Reverse-mode automatic differentiation and optimization of GPU kernels via enzymeWilliam S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim et al.SC 2021 · 50 citations
Related papers
- Iskra: A System for Inverse Geometry ProcessingAna Dodik, Ahmed H. Mahmoud, Justin SolomonSIGGRAPH 2026
- MeshFEM: A Block-accelerated Solver for Nonlinear Finite ElementsHaleh Mohammadian, Xinzhuo Hu, Roi Poranne, Julian PanettaSIGGRAPH 2026 · 1 citation
- YASPS: A Symbolic Framework for Extensible, High-Performance IPC SimulationXuan Tang, Kemeng Huang, Gilbert Bernstein, Minchen Li et al.SIGGRAPH 2026
- AD for an Array Language with Nested ParallelismRobert Schenck, Ola Rønning, Troels Henriksen, Cosmin E. OanceaSC 2022 · 11 citations
- Scalable Differentiable Physics for Learning and ControlYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2020 · 133 citations
