Gradients of Functions of Large Matrices
Nicholas Krämer, Pablo Moreno-Muñoz, Hrittik Roy, Søren Hauberg
Abstract
Tuning scientific and probabilistic machine learning models for example, partial differential equations, Gaussian processes, or Bayesian neural networks often relies on evaluating functions of matrices whose size grows with the data set or the number of parameters. While the state-of-the-art for evaluating these quantities is almost always based on Lanczos and Arnoldi iterations, the present work is the first to explain how to differentiate these workhorses of numerical linear algebra efficiently. To get there, we derive previously unknown adjoint systems for Lanczos and Arnoldi iterations, implement them in JAX, and show that the resulting code can compete with Diffrax when it comes to differentiating PDEs, GPyTorch for selecting Gaussian process models and beats standard factorisation methods for calibrating Bayesian neural networks. All this is achieved without any problem-specific code optimisation. Find the code at https://github.com/pnkraemer/experiments-lanczos-adjoints and install the library with pip install matfree.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7a6d35e-661d-4993-bf1b-d593c73e6527Builds on16
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- Bayesian Deep Learning via Subnetwork InferenceErik A. Daxberger, Eric T. Nalisnick, James Urquhart Allingham, Javier Antorán et al.ICML 2021 · 108 citations
Related papers
- CoLA: Exploiting Compositional Structure for Automatic and Efficient Numerical Linear AlgebraAndres Potapczynski, Marc Finzi, Geoff Pleiss, Andrew Gordon WilsonNeurIPS 2023 · 13 citations
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 1 citation
- Adjoint-aided inference of Gaussian process driven differential equationsPaterne Gahungu, Christopher W. Lanyon, Mauricio A. Álvarez, Engineer Bainomugisha et al.NeurIPS 2022 · 6 citations
- ΦFlow: Differentiable Simulations for PyTorch, TensorFlow and JaxPhilipp Holl, Nils ThuereyICML 2024 · 29 citations
- Hamiltonian Monte Carlo using an adjoint-differentiated Laplace approximation: Bayesian inference for latent Gaussian models and beyondCharles C. Margossian, Aki Vehtari, Daniel Simpson, Raj AgrawalNeurIPS 2020 · 30 citations
