MomentumRNN: Integrating Momentum into Recurrent Neural Networks
Tan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher, Bao Wang
Abstract
Designing deep neural networks is an art that often involves an expensive search over candidate architectures. To overcome this for recurrent neural nets (RNNs), we establish a connection between the hidden state dynamics in an RNN and gradient descent (GD). We then integrate momentum into this framework and propose a new family of RNNs, called MomentumRNNs. We theoretically prove and numerically demonstrate that MomentumRNNs alleviate the vanishing gradient issue in training RNNs. We study the momentum long-short term memory (MomentumLSTM) and verify its advantages in convergence speed and accuracy over its LSTM counterpart across a variety of benchmarks, with little compromise in computational or memory efficiency. We also demonstrate that MomentumRNN is applicable to many types of recurrent cells, including those in the state-of-the-art orthogonal RNNs. Finally, we show that other advanced momentum-based optimization methods, such as Adam and Nesterov accelerated gradients with a restart, can be easily incorporated into the MomentumRNN framework for designing new recurrent cells with even better performance. The code is available at this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen et al.NeurIPS 2021 · 75 citations
- Momentum Residual Neural NetworksMichael E. Sander, Pierre Ablin, Mathieu Blondel, Gabriel PeyréICML 2021 · 67 citations
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Lipschitz Recurrent Neural NetworksN. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson et al.ICLR 2021 · 32 citations
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher et al.NeurIPS 2022 · 28 citations
Builds on2
Related papers
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang et al.NeurIPS 2022 · 111 citations
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 25 citations
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 26 citations
- SBO-RNN: Reformulating Recurrent Neural Networks via Stochastic Bilevel OptimizationZiming Zhang, Yun Yue, Guojun Wu, Yanhua Li et al.NeurIPS 2021 · 4 citations
- Learning Low Dimensional State Spaces with Overparameterized Recurrent Neural NetsEdo Cohen-Karlik, Itamar Menuhin-Gruman, Raja Giryes, Nadav Cohen et al.ICLR 2023 · 2 citations
