High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to Algorithm
Jian Li, Yong Liu, Weiping Wang
摘要
Overparameterization often leads to benign overfitting, where deep neural networks can be trained to overfit the training data but still generalize well on unseen data. However, it lacks a generalized asymptotic framework for nonlinear regressions and connections to conventional complexity notions. In this paper, we propose a generalized high-dimensional analysis for nonlinear regression models, including various nonlinear feature mapping methods and subsampling. Specifically, we first provide an implicit regularization parameter and asymptotic equivalents related to a classical complexity notion, i.e., effective dimension. We then present a high-dimensional analysis for nonlinear ridge regression and extend it to ridgeless regression in the under-parameterized and over-parameterized regimes, respectively. We find that the limiting risks decrease with the effective dimension. Motivated by these theoretical findings, we propose an algorithm, namely RFRed, to improve generalization ability. Finally, we validate our theoretical findings and the proposed algorithm through several experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descentZhenyu Liao, Romain Couillet, Michael W. MahoneyNeurIPS 2020 · 被引用 102 次
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu 等ICLR 2020 · 被引用 77 次
- Sketched Ridgeless Linear Regression: The Role of DownsamplingXin Chen, Yicheng Zeng, Siyue Yang, Qiang SunICML 2023 · 被引用 8 次
- Optimal Convergence Rates for Agnostic Nyström Kernel LearningJian Li, Yong Liu, Weiping WangICML 2023 · 被引用 3 次
相关 Paper
- Theoretical Limitations of Ensembles in the Age of OverparameterizationNiclas Dern, John Patrick Cunningham, Geoff PleissICML 2025
- Interpolation can hurt robust generalization even when there is no noiseKonstantin Donhauser, Alexandru Tifrea, Michael Aerni, Reinhard Heckel 等NeurIPS 2021 · 被引用 18 次
- Optimal criterion for feature learning of two-layer linear neural network in high dimensional interpolation regimeKeita Suzuki, Taiji SuzukiICLR 2024 · 被引用 2 次
- Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methodsTaiji Suzuki, Shunta AkiyamaICLR 2021 · 被引用 12 次
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 被引用 9 次
