Lune

NeurIPS2022顶会

Generalization Bounds for Gradient Methods via Discrete and Continuous Prior

Xuanyuan Luo, Bei Luo, Jian Li

2022年份
5被引次数
3顶会引用

摘要

Proving algorithm-dependent generalization error bounds for gradient-type optimization methods has attracted significant attention recently in learning theory. However, most existing trajectory-based analyses require either restrictive assumptions on the learning rate (e.g., fast decreasing learning rate), or continuous injected noise (such as the Gaussian noise in Langevin dynamics). In this paper, we introduce a new discrete data-dependent prior to the PAC-Bayesian framework, and prove a high probability generalization bound of order O(1n⋅∑t=1T(γt/εt)2∥gt∥2)O(\frac{1}{n}\cdot \sum_{t=1}^T(\gamma_t/\varepsilon_t)^2\left\|{\mathbf{g}_t}\right\|^2) for Floored GD (i.e. a version of gradient descent with precision level εt\varepsilon_t), where nn is the number of training samples, γt\gamma_t is the learning rate at step tt, gt\mathbf{g}_t is roughly the difference of the gradient computed using all samples and that using only prior samples. ∥gt∥\left\|{\mathbf{g}_t}\right\| is upper bounded by and and typical much smaller than the gradient norm ∥∇f(Wt)∥\left\|{\nabla f(W_t)}\right\|. We remark that our bound holds for nonconvex and nonsmooth scenarios. Moreover, our theoretical results provide numerically favorable upper bounds of testing errors (e.g., 0.0370.037 on MNIST). Using a similar technique, we can also obtain new generalization bounds for certain variants of SGD. Furthermore, we study the generalization bounds for gradient Langevin Dynamics (GLD). Using the same framework with a carefully constructed continuous prior, we show a new high probability generalization bound of order O(1n+L2n2∑t=1T(γt/σt)2)O(\frac{1}{n} + \frac{L^2}{n^2}\sum_{t=1}^T(\gamma_t/\sigma_t)^2) for GLD. The new 1/n21/n^2 rate is due to the concentration of the difference between the gradient of training samples and that of the prior.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖