Momentum Residual Neural Networks
Michael E. Sander, Pierre Ablin, Mathieu Blondel, Gabriel Peyré
摘要
The training of deep residual neural networks (ResNets) with backpropagation has a memory cost that increases linearly with respect to the depth of the network. A way to circumvent this issue is to use reversible architectures. In this paper, we propose to change the forward rule of a ResNet by adding a momentum term. The resulting networks, momentum residual neural networks (Momentum ResNets), are invertible. Unlike previous invertible architectures, they can be used as a drop-in replacement for any existing ResNet block. We show that Momentum ResNets can be interpreted in the infinitesimal step size regime as second-order ordinary differential equations (ODEs) and exactly characterize how adding momentum progressively increases the representation capabilities of Momentum ResNets. Our analysis reveals that Momentum ResNets can learn any linear mapping up to a multiplicative factor, while ResNets cannot. In a learning to optimize setting, where convergence to a fixed point is required, we show theoretically and empirically that our method succeeds while existing invertible architectures fail. We show on CIFAR and ImageNet that Momentum ResNets have the same accuracy as ResNets, while having a much smaller memory footprint, and show that pre-trained Momentum ResNets are promising for fine-tuning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Efficient and Accurate Gradients for Neural SDEsPatrick Kidger, James Foster, Xuechen Li, Terry J. LyonsNeurIPS 2021 · 被引用 107 次
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen 等NeurIPS 2021 · 被引用 75 次
- A Dynamical System Perspective for Lipschitz Neural NetworksLaurent Meunier, Blaise Delattre, Alexandre Araujo, Alexandre AllauzenICML 2022 · 被引用 69 次
- Reversible Vision TransformersKarttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu 等CVPR 2022 · 被引用 52 次
- ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence GenerationBei Li, Quan Du, Tao Zhou, Yi Jing 等ACL 2022 · 被引用 43 次
它引用的顶会 Paper6
- Dissecting Neural ODEsStefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita 等NeurIPS 2020 · 被引用 261 次
- Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependenciesT. Konstantin Rusch, Siddhartha MishraICLR 2021 · 被引用 121 次
- On Second Order Behaviour in Augmented Neural ODEsAlexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski 等NeurIPS 2020 · 被引用 116 次
- Approximation Capabilities of Neural ODEs and Invertible Residual NetworksHan Zhang, Xi Gao, Jacob Unterman, Tom ArodzICML 2020 · 被引用 114 次
- MomentumRNN: Integrating Momentum into Recurrent Neural NetworksTan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher 等NeurIPS 2020 · 被引用 32 次
相关 Paper
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 被引用 42 次
- IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient ComputationYiting Chen, Zongwei Huo, Junchi YanICML 2026
- Memory-Efficient Reversible Spiking Neural NetworksHong Zhang, Yu ZhangAAAI 2024 · 被引用 13 次
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient FinetuningChen Zhao, Shuming Liu, Karttikeya Mangalam, Guocheng Qian 等CVPR 2024
