Lune

ICLR2022Top-tier venue

A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy Observations

Namjoon Suh, Hyunouk Ko, Xiaoming Huo

2022Year
15Citations
4Top-tier citations

Abstract

We study the generalization properties of the overparameterized deep neural network (DNN) with Rectified Linear Unit (ReLU) activations.Under the non-parametric regression framework, it is assumed that the ground-truth function is from a reproducing kernel Hilbert space (RKHS) induced by a neural tangent kernel (NTK) of ReLU DNN, and a dataset is given with the noises. Without a delicate adoption of early stopping, we prove that the overparametrized DNN trained by vanilla gradient descent does not recover the ground-truth function. It turns out that the estimated DNN's L2L_{2} prediction error is bounded away from 00. As a complement of the above result, we show that the ℓ2\ell_{2}-regularized gradient descent enables the overparametrized DNN achieve the minimax optimal convergence rate of the L2L_{2} prediction error, without early stopping. Notably, the rate we obtained is faster than O(n−1/2)\mathcal{O}(n^{-1/2}) known in the literature.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers4

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines