MomentumRNN: Integrating Momentum into Recurrent Neural Networks
Tan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher, Bao Wang
摘要
Designing deep neural networks is an art that often involves an expensive search over candidate architectures. To overcome this for recurrent neural nets (RNNs), we establish a connection between the hidden state dynamics in an RNN and gradient descent (GD). We then integrate momentum into this framework and propose a new family of RNNs, called MomentumRNNs. We theoretically prove and numerically demonstrate that MomentumRNNs alleviate the vanishing gradient issue in training RNNs. We study the momentum long-short term memory (MomentumLSTM) and verify its advantages in convergence speed and accuracy over its LSTM counterpart across a variety of benchmarks, with little compromise in computational or memory efficiency. We also demonstrate that MomentumRNN is applicable to many types of recurrent cells, including those in the state-of-the-art orthogonal RNNs. Finally, we show that other advanced momentum-based optimization methods, such as Adam and Nesterov accelerated gradients with a restart, can be easily incorporated into the MomentumRNN framework for designing new recurrent cells with even better performance. The code is available at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Heavy Ball Neural Ordinary Differential EquationsHedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen 等NeurIPS 2021 · 被引用 75 次
- Momentum Residual Neural NetworksMichael E. Sander, Pierre Ablin, Mathieu Blondel, Gabriel PeyréICML 2021 · 被引用 67 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
- Lipschitz Recurrent Neural NetworksN. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson 等ICLR 2021 · 被引用 32 次
- Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient MethodHo Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley J. Osher 等NeurIPS 2022 · 被引用 28 次
它引用的顶会 Paper2
相关 Paper
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang 等NeurIPS 2022 · 被引用 111 次
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 被引用 26 次
- SBO-RNN: Reformulating Recurrent Neural Networks via Stochastic Bilevel OptimizationZiming Zhang, Yun Yue, Guojun Wu, Yanhua Li 等NeurIPS 2021 · 被引用 4 次
- Learning Low Dimensional State Spaces with Overparameterized Recurrent Neural NetsEdo Cohen-Karlik, Itamar Menuhin-Gruman, Raja Giryes, Nadav Cohen 等ICLR 2023 · 被引用 2 次
