Understanding Generalization in Recurrent Neural Networks
Zhuozhuo Tu, Fengxiang He, Dacheng Tao
Abstract
In this work, we develop the theory for analyzing the generalization performance of recurrent neural networks. We first present a new generalization bound for recurrent neural networks based on matrix 1-norm and Fisher-Rao norm. The definition of Fisher-Rao norm relies on a structural lemma about the gradient of RNNs. This new generalization bound assumes that the covariance matrix of the input data is positive definite, which might limit its use in practice. To address this issue, we propose to add random noise to the input data and prove a generalization bound for training with random noise, which is an extension of the former one. Compared with existing results, our generalization bounds have no explicit dependency on the size of networks. We also discover that Fisher-Rao norm for RNNs can be interpreted as a measure of gradient, and incorporating this gradient measure not only can tighten the bound, but allows us to build a relationship between generalization and trainability. Based on the bound, we theoretically analyze the effect of covariance of features on generalization of RNNs and discuss how weight decay and gradient clipping in the training can help improve generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d04d6d9-cb66-4324-9408-9bdd3f03b94cCited by top-tier papers6
- Framing RNN as a kernel method: A neural ODE approachAdeline Fermanian, Pierre Marion, Jean-Philippe Vert, Gérard BiauNeurIPS 2021 · 34 citations
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes et al.NeurIPS 2024 · 14 citations
- From Generalization Analysis to Optimization Designs for State Space ModelsFusheng Liu, Qianxiao LiICML 2024 · 12 citations
- Learning and Generalization in RNNsAbhishek Panigrahi, Navin GoyalNeurIPS 2021 · 3 citations
- Nonparametric Quantile Regression with ReLU-Activated Recurrent Neural NetworksHan Yu, Lyumin Wu, Wenxin Zhou, Zhao RenNeurIPS 2025 · 1 citation
Related papers
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda et al.ICLR 2024 · 6 citations
- On the Provable Generalization of Recurrent Neural NetworksLifu Wang, Bo Shen, Bo Hu, Xing CaoNeurIPS 2021 · 9 citations
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- Network-to-Network Regularization: Enforcing Occam's Razor to Improve GeneralizationRohan Ghosh, Mehul MotaniNeurIPS 2021 · 6 citations
- PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNsDeividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski et al.AAAI 2024 · 5 citations
