Tuning the burn-in phase in training recurrent neural networks improves their performance
Julian D. Schiller, Malte Heinrich, Victor G. Lopez, Matthias A. Müller
Abstract
Training recurrent neural networks (RNNs) with standard backpropagation through time (BPTT) can be challenging, especially in the presence of long input sequences. A practical alternative to reduce computational and memory overhead is to perform BPTT repeatedly over shorter segments of the training data set, corresponding to truncated BPTT. In this paper, we examine the training of RNNs when using such a truncated learning approach for time series tasks. Specifically, we establish theoretical bounds on the accuracy and performance loss when optimizing over subsequences instead of the full data sequence. This reveals that the burn-in phase of the RNN is an important tuning knob in its training, with significant impact on the performance guarantees. We validate our theoretical results through experiments on standard benchmarks from the fields of system identification and time series forecasting. In all experiments, we observe a strong influence of the burn-in phase on the training process, and proper tuning can lead to a reduction of the prediction error on the training and test data of more than 60% in some cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1da536e1-3e07-4f83-9758-b0bf29bc69c3Builds on8
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 38 citations
- RNNs of RNNs: Recursive Construction of Stable Assemblies of Recurrent Neural NetworksLeo Kozachkov, Michaela Ennis, Jean-Jacques E. SlotineNeurIPS 2022 · 30 citations
- Improved Worst-Case Regret Bounds for Randomized Least-Squares Value IterationPriyank Agrawal, Jinglin Chen, Nan JiangAAAI 2021 · 24 citations
Related papers
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 51 citations
- Training Recurrent Neural Networks via Forward Propagation Through TimeAnil Kag, Venkatesh SaligramaICML 2021 · 48 citations
- SkipW: Resource Adaptable RNN with Strict Upper Computational LimitTsiry Mayet, Anne Lambert, Pascal Leguyadec, Françoise Le Bolzer et al.ICLR 2021
- Training Recurrent Neural Networks Online by Learning Explicit State VariablesSomjit Nath, Vincent Liu, Alan Chan, Xin Li et al.ICLR 2020 · 9 citations
- Balanced Resonate-and-Fire NeuronsSaya Higuchi, Sebastian Kairat, Sander M. Bohté, Sebastian OtteICML 2024 · 19 citations
