Gradients of Functions of Large Matrices
Nicholas Krämer, Pablo Moreno-Muñoz, Hrittik Roy, Søren Hauberg
摘要
Tuning scientific and probabilistic machine learning models for example, partial differential equations, Gaussian processes, or Bayesian neural networks often relies on evaluating functions of matrices whose size grows with the data set or the number of parameters. While the state-of-the-art for evaluating these quantities is almost always based on Lanczos and Arnoldi iterations, the present work is the first to explain how to differentiate these workhorses of numerical linear algebra efficiently. To get there, we derive previously unknown adjoint systems for Lanczos and Arnoldi iterations, implement them in JAX, and show that the resulting code can compete with Diffrax when it comes to differentiating PDEs, GPyTorch for selecting Gaussian process models and beats standard factorisation methods for calibrating Bayesian neural networks. All this is achieved without any problem-specific code optimisation. Find the code at https://github.com/pnkraemer/experiments-lanczos-adjoints and install the library with pip install matfree.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 被引用 344 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
- Bayesian Deep Learning via Subnetwork InferenceErik A. Daxberger, Eric T. Nalisnick, James Urquhart Allingham, Javier Antorán 等ICML 2021 · 被引用 108 次
相关 Paper
- CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear AlgebraAndres Potapczynski, Marc Finzi, Geoff Pleiss, Andrew Gordon WilsonNeurIPS 2023 · 被引用 13 次
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 被引用 1 次
- Adjoint-aided inference of Gaussian process driven differential equationsPaterne Gahungu, Christopher W. Lanyon, Mauricio A. Álvarez, Engineer Bainomugisha 等NeurIPS 2022 · 被引用 6 次
- ΦFlow: Differentiable Simulations for PyTorch, TensorFlow and JaxPhilipp Holl, Nils ThuereyICML 2024 · 被引用 29 次
- Hamiltonian Monte Carlo using an adjoint-differentiated Laplace approximation: Bayesian inference for latent Gaussian models and beyondCharles C. Margossian, Aki Vehtari, Daniel Simpson, Raj AgrawalNeurIPS 2020 · 被引用 30 次
