Fast Saturating Gate for Learning Long Time Scales with Recurrent Neural Networks
Kentaro Ohno, Sekitoshi Kanai, Yasutoshi Ida
Abstract
Gate functions in recurrent models, such as an LSTM and GRU, play a central role in learning various time scales in modeling time series data by using a bounded activation function. However, it is difficult to train gates to capture extremely long time scales due to gradient vanishing of the bounded function for large inputs, which is known as the saturation problem. We closely analyze the relation between saturation of the gate function and efficiency of the training. We prove that the gradient vanishing of the gate function can be mitigated by accelerating the convergence of the saturating function, i.e., making the output of the function converge to 0 or 1 faster. Based on the analysis results, we propose a gate function called fast gate that has a doubly exponential convergence rate with respect to inputs by simple function composition. We empirically show that our method outperforms previous methods in accuracy and computational efficiency on benchmark tasks involving extremely long time scales.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b32e654-fbeb-45ad-b97d-6aa2f7ec5591Builds on6
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- CKConv: Continuous Kernel Convolution For Sequential DataDavid W. Romero, Anna Kuzina, Erik J. Bekkers, Jakub Mikolaj Tomczak et al.ICLR 2022 · 149 citations
- Improving the Gating Mechanism of Recurrent Neural NetworksAlbert Gu, Çaglar Gülçehre, Thomas Paine, Matt Hoffman et al.ICML 2020 · 111 citations
- UnICORNN: A recurrent model for learning very long time dependenciesT. Konstantin Rusch, Siddhartha MishraICML 2021 · 76 citations
- Lipschitz Recurrent Neural NetworksN. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson et al.ICLR 2021 · 32 citations
Related papers
- Time Adaptive Recurrent Neural NetworkAnil Kag, Venkatesh SaligramaCVPR 2021
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 33 citations
- Long Expressive Memory for Sequence ModelingT. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. MahoneyICLR 2022 · 57 citations
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 51 citations
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 109 citations
