Understanding Generalization in Recurrent Neural Networks
Zhuozhuo Tu, Fengxiang He, Dacheng Tao
摘要
In this work, we develop the theory for analyzing the generalization performance of recurrent neural networks. We first present a new generalization bound for recurrent neural networks based on matrix 1-norm and Fisher-Rao norm. The definition of Fisher-Rao norm relies on a structural lemma about the gradient of RNNs. This new generalization bound assumes that the covariance matrix of the input data is positive definite, which might limit its use in practice. To address this issue, we propose to add random noise to the input data and prove a generalization bound for training with random noise, which is an extension of the former one. Compared with existing results, our generalization bounds have no explicit dependency on the size of networks. We also discover that Fisher-Rao norm for RNNs can be interpreted as a measure of gradient, and incorporating this gradient measure not only can tighten the bound, but allows us to build a relationship between generalization and trainability. Based on the bound, we theoretically analyze the effect of covariance of features on generalization of RNNs and discuss how weight decay and gradient clipping in the training can help improve generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Framing RNN as a kernel method: A neural ODE approachAdeline Fermanian, Pierre Marion, Jean-Philippe Vert, Gérard BiauNeurIPS 2021 · 被引用 34 次
- Provable Benefits of Complex Parameterizations for Structured State Space ModelsYuval Ran-Milo, Eden Lumbroso, Edo Cohen-Karlik, Raja Giryes 等NeurIPS 2024 · 被引用 14 次
- From Generalization Analysis to Optimization Designs for State Space ModelsFusheng Liu, Qianxiao LiICML 2024 · 被引用 12 次
- Learning and Generalization in RNNsAbhishek Panigrahi, Navin GoyalNeurIPS 2021 · 被引用 3 次
- Nonparametric Quantile Regression with ReLU-Activated Recurrent Neural NetworksHan Yu, Lyumin Wu, Wenxin Zhou, Zhao RenNeurIPS 2025 · 被引用 1 次
相关 Paper
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda 等ICLR 2024 · 被引用 6 次
- On the Provable Generalization of Recurrent Neural NetworksLifu Wang, Bo Shen, Bo Hu, Xing CaoNeurIPS 2021 · 被引用 9 次
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
- Network-to-Network Regularization: Enforcing Occam's Razor to Improve GeneralizationRohan Ghosh, Mehul MotaniNeurIPS 2021 · 被引用 6 次
- PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNsDeividas Eringis, John Leth, Zheng-Hua Tan, Rafael Wisniewski 等AAAI 2024 · 被引用 5 次
