Randomized Automatic Differentiation
Deniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson, Ryan P. Adams
Abstract
The successes of deep learning, variational inference, and many other fields have been aided by specialized implementations of reverse-mode automatic differentiation (AD) to compute gradients of mega-dimensional objectives. The AD techniques underlying these tools were designed to compute exact gradients to numerical precision, but modern machine learning models are almost always trained with stochastic gradient descent. Why spend computation and memory on exact (minibatch) gradients only to use them for stochastic optimization? We develop a general framework and approach for randomized automatic differentiation (RAD), which can allow unbiased gradient estimates to be computed with reduced memory in return for variance. We examine limitations of the general approach, and argue that we must leverage problem specific structure to realize benefits. We develop RAD techniques for a variety of simple neural network architectures, and show that for a fixed memory budget, RAD converges in fewer iterations than using a small batch size for feedforward networks, and in a similar number for recurrent networks. We also show that RAD can be applied to scientific computing, and use it to develop a low-memory stochastic gradient method for optimizing the control parameters of a linear reaction-diffusion PDE representing a fission reactor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d7b5e67-c1fa-4270-bfb0-c16039aa0da8Cited by top-tier papers15
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 113 citations
- Rectangular Flows for Manifold LearningAnthony L. Caterini, Gabriel Loaiza-Ganem, Geoff Pleiss, John P. CunninghamNeurIPS 2021 · 58 citations
- Dataset Distillation with Convexified Implicit GradientsNoel Loo, Ramin M. Hasani, Mathias Lechner, Daniela RusICML 2023 · 56 citations
- L2ight: Enabling On-Chip Learning for Optical Neural Networks via Efficient in-situ Subspace OptimizationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Zixuan Jiang et al.NeurIPS 2021 · 41 citations
Builds on2
- Faster Neural Network Training with Approximate Tensor OperationsMenachem Adelman, Kfir Y. Levy, Ido Hakimi, Mark SilbersteinNeurIPS 2021 · 30 citations
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu et al.ICLR 2020 · 29 citations
Related papers
- Storchastic: A Framework for General Stochastic Automatic DifferentiationEmile van Krieken, Jakub M. Tomczak, Annette ten TeijeNeurIPS 2021 · 19 citations
- Automatic Differentiation of Programs with Discrete RandomnessGaurav Arya, Moritz Schauer, Frank Schäfer, Christopher RackauckasNeurIPS 2022 · 56 citations
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 1 citation
- Stochastic Taylor Derivative Estimator: Efficient amortization for arbitrary differential operatorsZekun Shi, Zheyuan Hu, Min Lin, Kenji KawaguchiNeurIPS 2024 · 32 citations
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver HeuristicsAvik Pal, Yingbo Ma, Viral B. Shah, Christopher Vincent RackauckasICML 2021 · 44 citations
