Lune

NeurIPS2021顶会

On the Provable Generalization of Recurrent Neural Networks

Lifu Wang, Bo Shen, Bo Hu, Xing Cao

2021年份
9被引次数
2顶会引用

摘要

Recurrent Neural Network (RNN) is a fundamental structure in deep learning. Recently, some works study the training process of over-parameterized neural networks, and show that over-parameterized networks can learn functions in some notable concept classes with a provable generalization error bound. In this paper, we analyze the training and generalization for RNNs with random initialization, and provide the following improvements over recent works: 1) For a RNN with input sequence x=(X1,X2,...,XL)x=(X_1,X_2,...,X_L), previous works study to learn functions that are summation of f(βlTXl)f(\beta^T_lX_l) and require normalized conditions that ∣∣Xl∣∣≤ϵ||X_l||\leq\epsilon with some very small ϵ\epsilon depending on the complexity of ff. In this paper, using detailed analysis about the neural tangent kernel matrix, we prove a generalization error bound to learn such functions without normalized conditions and show that some notable concept classes are learnable with the numbers of iterations and samples scaling almost-polynomially in the input length LL. 2) Moreover, we prove a novel result to learn N-variables functions of input sequence with the form f(βT[Xl1,...,XlN])f(\beta^T[X_{l_1},...,X_{l_N}]), which do not belong to the"additive"concept class, i,e., the summation of function f(Xl)f(X_l). And we show that when either NN or l0=max⁡(l1,..,lN)−min⁡(l1,..,lN)l_0=\max(l_1,..,l_N)-\min(l_1,..,l_N) is small, f(βT[Xl1,...,XlN])f(\beta^T[X_{l_1},...,X_{l_N}]) will be learnable with the number iterations and samples scaling almost-polynomially in the input length LL.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 08c1e789-e986-4371-8009-29c834d69b3f

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖