A Kernel Perspective of Skip Connections in Convolutional Networks
Daniel Barzilai, Amnon Geifman, Meirav Galun, Ronen Basri
Abstract
Over-parameterized residual networks are amongst the most successful convolutional neural architectures for image processing. Here we study their properties through their Gaussian Process and Neural Tangent kernels. We derive explicit formulas for these kernels, analyze their spectra and provide bounds on their implied condition numbers. Our results indicate that (1) with ReLU activation, the eigenvalues of these residual kernels decay polynomially at a similar rate as the same kernels when skip connections are not used, thus maintaining a similar frequency bias; (2) however, residual kernels are more locally biased. Our analysis further shows that the matrices obtained by these residual kernels yield favorable condition numbers at finite depths than those obtained without the skip connections, enabling therefore faster convergence of training with gradient descent. Published as a conference paper at ICLR 2023 work of Lee et al. (2019); Xiao et al. (2020) and Chen et al. (2021), who related between the condition number of NTK and the trainability of corresponding finite width networks. 3 PRELIMINARIES We consider mutli-channel 1-D input signals x ∈ R C0×d of length d with C 0 channels. We use 1-D input signals to simplify notations and note that all our results can naturally be extended to 2-D signals. Let MS (C 0 , d) = S C0-1 × . . . × S C0-1 d times ⊆ √ d S dC0-1 be the multi-sphere, so x = (x 1 , ..., x d ) ∈ MS (C 0 , d) iff ∀i ∈ [d], x i = 1. For our analysis, we assume that the input signals are distributed uniformly on the multi-sphere. The discrete convolution of a filter w ∈ R q with a vector v ∈ R d is defined as where 1 ≤ i ≤ d. We use circular padding, so indices [v] j with j ≤ 0 and j > d are well defined. We use multi-index notation denoted by bold letters, i.e., n, k ∈ N d , where N is the set of natural numbers including zero. b n , λ k ∈ R are scalars that depend on n, k, and for t ∈ R d we let t n = t n1 1 • ... • t n d d . As is convention, we say that n ≥ k iff n i ≥ k i for all i ∈ [d]. Thus, the power series n≥0 b n t n should read n1≥0,n2≥0,... b n1,n2,... t n1 1 t n2 2 ... We further use the following notation to denote sub-vectors and sub-matrices. ∀i ∈ N, let D that for a matrix M we can write: 2 . We use (s i v) j = v j+i to denote the cyclic shift of v to the left by i pixels. Finally, for every kernel K : R d × R d → R we define the normalized kernel to be K (x, z) = K(x,z) √ K(x,x)K(z,z)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generalization in Kernel Regression Under Realistic AssumptionsDaniel Barzilai, Ohad ShamirICML 2024 · 22 citations
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin et al.AAAI 2024 · 7 citations
- Generalization of Scaled Deep ResNets in the Mean-Field RegimeYihang Chen, Fanghui Liu, Yiping Lu, Grigorios Chrysos et al.ICLR 2024 · 2 citations
- On the Spectral Differences Between NTK and CNTK and Their Implications for Point Cloud RecognitionYuanqu Mou, Chang Gou, Haiyang Bai, Jia LiuICLR 2026
Builds on13
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs et al.ICML 2020 · 229 citations
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov et al.ICLR 2020 · 167 citations
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun et al.NeurIPS 2020 · 118 citations
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 107 citations
Related papers
- On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process KernelsAmnon Geifman, Meirav Galun, David Jacobs, Ronen BasriNeurIPS 2022 · 22 citations
- On the Random Conjugate Kernel and Neural Tangent KernelZhengmian Hu, Heng HuangICML 2021 · 15 citations
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 23 citations
- Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained AnalysisWuyang Chen, Wei Huang, Xinyu Gong, Boris Hanin et al.NeurIPS 2022 · 9 citations
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 14 citations
