Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
Kaiqi Zhang, Yu-Xiang Wang
Abstract
We study the theory of neural network (NN) from the lens of classical nonparametric regression problems with a focus on NN's ability to adaptively estimate functions with heterogeneous smoothness -a property of functions in Besov or Bounded Variation (BV) classes. Existing work on this problem requires tuning the NN architecture based on the function spaces and sample size. We consider a "Parallel NN" variant of deep ReLU networks and show that the standard ℓ 2 regularization is equivalent to promoting the ℓ p -sparsity (0 < p < 1) in the coefficient vector of an end-to-end learned function bases, i.e., a dictionary. Using this equivalence, we further establish that by tuning only the regularization factor, such parallel NN achieves an estimation error arbitrarily close to the minimax rates for both the Besov and BV classes. Notably, it gets exponentially closer to minimax optimal as the NN gets deeper. Our research sheds new lights on why depth matters and how NNs are more powerful than kernel methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d91ceb11-b7f5-42a7-a23c-6c2ffed273d0Cited by top-tier papers7
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry et al.NeurIPS 2024 · 15 citations
- Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering PhenomenonTongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul ParhiNeurIPS 2025 · 8 citations
- Nonparametric Classification on Low Dimensional Manifolds using Overparameterized Convolutional Residual NetworksZixuan Zhang, Kaiqi Zhang, Minshuo Chen, Yuma Takeda et al.NeurIPS 2024 · 6 citations
- How many samples are needed to train a deep neural network?Pegah Golestaneh, Mahsa Taheri, Johannes LedererICLR 2025
- Sharp Generalization for Nonparametric Regression by Over-Parameterized Neural Networks: A Distribution-Free Analysis in Spherical CovariateYingzhen YangICML 2025
Builds on10
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer NetworksMert Pilanci, Tolga ErgenICML 2020 · 142 citations
- Revealing the Structure of Deep Neural Networks via Convex DualityTolga Ergen, Mert PilanciICML 2021 · 77 citations
Related papers
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
- Convergence Rates of Variational Inference in Sparse Deep LearningBadr-Eddine Chérief-AbdellatifICML 2020 · 43 citations
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 15 citations
- Optimality and Adaptivity of Deep Neural Features for Instrumental Variable RegressionJuno Kim, Dimitri Meunier, Arthur Gretton, Taiji Suzuki et al.ICLR 2025
- Deep Neural Network Regression with Functional CovariatesHang Zhou, Ju-Sheng Hong, Xiucai Ding, Jane-Ling WangICML 2026
