ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation Guarantees
Kuan-Lin Chen, Ching Hua Lee, Harinath Garudadri, Bhaskar D. Rao
摘要
Models recently used in the literature proving residual networks (ResNets) are better than linear predictors are actually different from standard ResNets that have been widely used in computer vision. In addition to the assumptions such as scalar-valued output or single residual block, the models fundamentally considered in the literature have no nonlinearities at the final residual representation that feeds into the final affine layer. To codify such a difference in nonlinearities and reveal a linear estimation property, we define ResNEsts, i.e., Residual Nonlinear Estimators, by simply dropping nonlinearities at the last residual representation from standard ResNets. We show that wide ResNEsts with bottleneck blocks can always guarantee a very desirable training property that standard ResNets aim to achieve, i.e., adding more blocks does not decrease performance given the same set of basis elements. To prove that, we first recognize ResNEsts are basis function models that are limited by a coupling problem in basis learning and linear prediction. Then, to decouple prediction weights from basis learning, we construct a special architecture termed augmented ResNEst (A-ResNEst) that always guarantees no worse performance with the addition of a block. As a result, such an A-ResNEst establishes empirical risk lower bounds for a ResNEst using corresponding bases. Our results demonstrate ResNEsts indeed have a problem of diminishing feature reuse; however, it can be avoided by sufficiently expanding or widening the input space, leading to the above-mentioned desirable property. Inspired by the densely connected networks (DenseNets) that have been shown to outperform ResNets, we also propose a corresponding new model called Densely connected Nonlinear Estimator (DenseNEst). We show that any DenseNEst can be represented as a wide ResNEst with bottleneck blocks. Unlike ResNEsts, DenseNEsts exhibit the desirable property without any special architectural re-design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Deep Residual-Dense Lattice Network for Speech EnhancementMohammad Nikzad, Aaron Nicolson, Yongsheng Gao, Jun Zhou 等AAAI 2020 · 被引用 42 次
- Characterizing ResNet's Universal Approximation CapabilityChenghao Liu, Enming Liang, Minghua ChenICML 2024
- Residual Alignment: Uncovering the Mechanisms of Residual NetworksJianing Li, Vardan PapyanNeurIPS 2023 · 被引用 21 次
- Nonparametric Classification on Low Dimensional Manifolds using Overparameterized Convolutional Residual NetworksZixuan Zhang, Kaiqi Zhang, Minshuo Chen, Yuma Takeda 等NeurIPS 2024 · 被引用 6 次
- Approximation theory for 1-Lipschitz ResNetsDavide Murari, Takashi Furuya, Carola-Bibiane SchönliebNeurIPS 2025 · 被引用 7 次
