Locality-Aware Automatic Differentiation on the GPU for Mesh-Based Computations
Ahmed H. Mahmoud, Rahul Goel, Jonathan Ragan-Kelley, Justin Solomon
摘要
Fig. 1. We introduce a GPU system for efficient automatic differentiation of computations defined on triangle meshes that exploits locality and sparsity in mesh-based workloads. Using our system, users specify only the energy terms of their application while our system computes gradients, sparse Hessians, and Jacobians automatically and efficiently on the GPU. Here, we use a Newton solver for large-scale elastic shell simulation to simulate ≈ 700 Spot cows (≈2.1M vertices in total) falling to the ground where derivative computation accounts for only 12.2% of the total runtime.
We present a GPU-based system for automatic differentiation (AD) of functions defined on triangle meshes, designed to exploit the locality and sparsity in mesh-based computation. Our system evaluates derivatives using perelement forward-mode AD, confining all computation to registers and shared memory and assembling global gradients, sparse Jacobians, and sparse Hessians directly on the GPU. By avoiding global computation graphs, intermediate buffers, and device-host synchronization, our approach minimizes memory traffic and enables efficient differentiation under both static and dynamically changing sparsity. Our programming model lets users express energy terms over mesh neighborhoods, while our system automatically manages parallel execution, derivative propagation, sparse assembly, and matrix-free operations such as Hessian-vector products. Our system supports both scalar-and vector-valued objectives, dynamic interaction-driven sparsity updates, and seamless integration with external GPU sparse linear solvers. We evaluate our system on applications including elastic and cloth simulation, surface parameterization, mesh smoothing, frame field design, ARAP deformation, and spherical manifold optimization. Across these tasks, our system consistently outperforms state-of-the-art differentiation frameworks, including PyTorch, JAX, Warp, Dr.JIT, EnzymeAD, and Thallo. We
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- Incremental potential contact: intersection-and inversion-free, large-deformation dynamicsMinchen Li, Zachary Ferguson, Teseo Schneider, Timothy R. Langlois 等SIGGRAPH 2020 · 被引用 320 次
- DR.JIT: a just-in-time compiler for differentiable renderingWenzel Jakob, Sébastien Speierer, Nicolas Roussel, Delio ViciniSIGGRAPH 2022 · 被引用 160 次
- Reverse-mode automatic differentiation and optimization of GPU kernels via enzymeWilliam S. Moses, Valentin Churavy, Ludger Paehler, Jan Hückelheim 等SC 2021 · 被引用 50 次
相关 Paper
- Iskra: A System for Inverse Geometry ProcessingAna Dodik, Ahmed H. Mahmoud, Justin SolomonSIGGRAPH 2026
- MeshFEM: A Block-accelerated Solver for Nonlinear Finite ElementsHaleh Mohammadian, Xinzhuo Hu, Roi Poranne, Julian PanettaSIGGRAPH 2026 · 被引用 1 次
- YASPS: A Symbolic Framework for Extensible, High-Performance IPC SimulationXuan Tang, Kemeng Huang, Gilbert Bernstein, Minchen Li 等SIGGRAPH 2026
- AD for an Array Language with Nested ParallelismRobert Schenck, Ola Rønning, Troels Henriksen, Cosmin E. OanceaSC 2022 · 被引用 11 次
- Scalable Differentiable Physics for Learning and ControlYi-Ling Qiao, Junbang Liang, Vladlen Koltun, Ming C. LinICML 2020 · 被引用 133 次
