On the Random Conjugate Kernel and Neural Tangent Kernel
Zhengmian Hu, Heng Huang
摘要
We investigate the distributions of Conjugate Kernel (CK) and Neural Tangent Kernel (NTK) for ReLU networks with random initialization. We derive the precise distributions and moments of the diagonal elements of these kernels. For a feedforward network, these values converge in law to a log-normal distribution when the network depth d and width n simultaneously tend to infinity and the variance of log diagonal elements is proportional to d/n. For the residual network, in the limit that number of branches m increases to infinity and the width n remains fixed, the diagonal elements of Conjugate Kernel converge in law to a log-normal distribution where the variance of log value is proportional to 1/n, and the diagonal elements of NTK converge in law to a log-normal distributed variable times the conjugate kernel of one feedforward network. Our new theoretical analysis results suggest that residual network remains trainable in the limit of infinite branches and fixed network width. The numerical experiments are conducted and all results validate the soundness of our theoretical analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Shaped Transformer: Attention Models in the Infinite Depth-and-Width LimitLorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He 等NeurIPS 2023 · 被引用 59 次
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 被引用 51 次
- Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and InitializationMariia Seleznova, Gitta KutyniokICML 2022 · 被引用 34 次
- Thompson Sampling in Function Spaces via Neural OperatorsRafael Oliveira, Xuesong Wang, Kian Ming A. Chai, Edwin V. BonillaNeurIPS 2025 · 被引用 2 次
- Optimization and Bayes: A Trade-off for Overparameterized Neural NetworksZhengmian Hu, Heng HuangNeurIPS 2023 · 被引用 1 次
它引用的顶会 Paper6
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 被引用 169 次
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 被引用 127 次
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 被引用 107 次
- Deep Networks and the Multiple Manifold ProblemSam Buchanan, Dar Gilboa, John WrightICLR 2021 · 被引用 9 次
相关 Paper
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 被引用 101 次
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 被引用 1 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- Divergence of Neural Tangent Kernel in Classification ProblemsZixiong Yu, Songtao Tian, Guhan ChenICLR 2025
- Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural NetworksChaoyue Liu, Han Bi, Like Hui, Xiao LiuNeurIPS 2025
