A Kernel Perspective of Skip Connections in Convolutional Networks
Daniel Barzilai, Amnon Geifman, Meirav Galun, Ronen Basri
摘要
Over-parameterized residual networks are amongst the most successful convolutional neural architectures for image processing. Here we study their properties through their Gaussian Process and Neural Tangent kernels. We derive explicit formulas for these kernels, analyze their spectra and provide bounds on their implied condition numbers. Our results indicate that (1) with ReLU activation, the eigenvalues of these residual kernels decay polynomially at a similar rate as the same kernels when skip connections are not used, thus maintaining a similar frequency bias; (2) however, residual kernels are more locally biased. Our analysis further shows that the matrices obtained by these residual kernels yield favorable condition numbers at finite depths than those obtained without the skip connections, enabling therefore faster convergence of training with gradient descent. Published as a conference paper at ICLR 2023 work of Lee et al. (2019); Xiao et al. (2020) and Chen et al. (2021), who related between the condition number of NTK and the trainability of corresponding finite width networks. 3 PRELIMINARIES We consider mutli-channel 1-D input signals x ∈ R C0×d of length d with C 0 channels. We use 1-D input signals to simplify notations and note that all our results can naturally be extended to 2-D signals. Let MS (C 0 , d) = S C0-1 × . . . × S C0-1 d times ⊆ √ d S dC0-1 be the multi-sphere, so x = (x 1 , ..., x d ) ∈ MS (C 0 , d) iff ∀i ∈ [d], x i = 1. For our analysis, we assume that the input signals are distributed uniformly on the multi-sphere. The discrete convolution of a filter w ∈ R q with a vector v ∈ R d is defined as where 1 ≤ i ≤ d. We use circular padding, so indices [v] j with j ≤ 0 and j > d are well defined. We use multi-index notation denoted by bold letters, i.e., n, k ∈ N d , where N is the set of natural numbers including zero. b n , λ k ∈ R are scalars that depend on n, k, and for t ∈ R d we let t n = t n1 1 • ... • t n d d . As is convention, we say that n ≥ k iff n i ≥ k i for all i ∈ [d]. Thus, the power series n≥0 b n t n should read n1≥0,n2≥0,... b n1,n2,... t n1 1 t n2 2 ... We further use the following notation to denote sub-vectors and sub-matrices. ∀i ∈ N, let D that for a matrix M we can write: 2 . We use (s i v) j = v j+i to denote the cyclic shift of v to the left by i pixels. Finally, for every kernel K : R d × R d → R we define the normalized kernel to be K (x, z) = K(x,z) √ K(x,x)K(z,z)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generalization in Kernel Regression Under Realistic AssumptionsDaniel Barzilai, Ohad ShamirICML 2024 · 被引用 22 次
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin 等AAAI 2024 · 被引用 7 次
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos 等ICLR 2024 · 被引用 2 次
- On the Spectral Differences Between NTK and CNTK and Their Implications for Point Cloud RecognitionYuanqu Mou, Chang Gou, Haiyang Bai, Jia LiuICLR 2026
它引用的顶会 Paper13
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs 等ICML 2020 · 被引用 229 次
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun 等NeurIPS 2020 · 被引用 118 次
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 被引用 107 次
相关 Paper
- On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process KernelsAmnon Geifman, Meirav Galun, David Jacobs, Ronen BasriNeurIPS 2022 · 被引用 22 次
- On the Random Conjugate Kernel and Neural Tangent KernelZhengmian Hu, Heng HuangICML 2021 · 被引用 15 次
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 被引用 23 次
- Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained AnalysisWuyang Chen, Wei Huang, Xinyu Gong, Boris Hanin 等NeurIPS 2022 · 被引用 9 次
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 被引用 14 次
