Randomized Automatic Differentiation
Deniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson, Ryan P. Adams
摘要
The successes of deep learning, variational inference, and many other fields have been aided by specialized implementations of reverse-mode automatic differentiation (AD) to compute gradients of mega-dimensional objectives. The AD techniques underlying these tools were designed to compute exact gradients to numerical precision, but modern machine learning models are almost always trained with stochastic gradient descent. Why spend computation and memory on exact (minibatch) gradients only to use them for stochastic optimization? We develop a general framework and approach for randomized automatic differentiation (RAD), which can allow unbiased gradient estimates to be computed with reduced memory in return for variance. We examine limitations of the general approach, and argue that we must leverage problem specific structure to realize benefits. We develop RAD techniques for a variety of simple neural network architectures, and show that for a fixed memory budget, RAD converges in fewer iterations than using a small batch size for feedforward networks, and in a similar number for recurrent networks. We also show that RAD can be applied to scientific computing, and use it to develop a low-memory stochastic gradient method for optimizing the control parameters of a linear reaction-diffusion PDE representing a fission reactor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- Flora: Low-Rank Adapters Are Secretly Gradient CompressorsYongchang Hao, Yanshuai Cao, Lili MouICML 2024 · 被引用 113 次
- Rectangular Flows for Manifold LearningAnthony L. Caterini, Gabriel Loaiza-Ganem, Geoff Pleiss, John P. CunninghamNeurIPS 2021 · 被引用 58 次
- Dataset Distillation with Convexified Implicit GradientsNoel Loo, Ramin M. Hasani, Mathias Lechner, Daniela RusICML 2023 · 被引用 56 次
- L2ight: Enabling On-Chip Learning for Optical Neural Networks via Efficient in-situ Subspace OptimizationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Zixuan Jiang 等NeurIPS 2021 · 被引用 41 次
它引用的顶会 Paper2
- Faster Neural Network Training with Approximate Tensor OperationsMenachem Adelman, Kfir Y. Levy, Ido Hakimi, Mark SilbersteinNeurIPS 2021 · 被引用 30 次
- SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable ModelsYucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu 等ICLR 2020 · 被引用 29 次
相关 Paper
- Storchastic: A Framework for General Stochastic Automatic DifferentiationEmile van Krieken, Jakub M. Tomczak, Annette ten TeijeNeurIPS 2021 · 被引用 19 次
- Automatic Differentiation of Programs with Discrete RandomnessGaurav Arya, Moritz Schauer, Frank Schäfer, Christopher RackauckasNeurIPS 2022 · 被引用 56 次
- Collapsing Taylor Mode Automatic DifferentiationFelix Dangel, Tim Siebert, Marius Zeinhofer, Andrea WaltherNeurIPS 2025 · 被引用 1 次
- Stochastic Taylor Derivative Estimator: Efficient amortization for arbitrary differential operatorsZekun Shi, Zheyuan Hu, Min Lin, Kenji KawaguchiNeurIPS 2024 · 被引用 32 次
- Opening the Blackbox: Accelerating Neural Differential Equations by Regularizing Internal Solver HeuristicsAvik Pal, Yingbo Ma, Viral B. Shah, Christopher Vincent RackauckasICML 2021 · 被引用 44 次
