On the Random Conjugate Kernel and Neural Tangent Kernel
Zhengmian Hu, Heng Huang
Abstract
We investigate the distributions of Conjugate Kernel (CK) and Neural Tangent Kernel (NTK) for ReLU networks with random initialization. We derive the precise distributions and moments of the diagonal elements of these kernels. For a feedforward network, these values converge in law to a log-normal distribution when the network depth d and width n simultaneously tend to infinity and the variance of log diagonal elements is proportional to d/n. For the residual network, in the limit that number of branches m increases to infinity and the width n remains fixed, the diagonal elements of Conjugate Kernel converge in law to a log-normal distribution where the variance of log value is proportional to 1/n, and the diagonal elements of NTK converge in law to a log-normal distributed variable times the conjugate kernel of one feedforward network. Our new theoretical analysis results suggest that residual network remains trainable in the limit of infinite branches and fixed network width. The numerical experiments are conducted and all results validate the soundness of our theoretical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa7fdc25-386f-44bf-b121-7abf80656e95Cited by top-tier papers7
- The Shaped Transformer: Attention Models in the Infinite Depth-and-Width LimitLorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He et al.NeurIPS 2023 · 59 citations
- The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at InitializationMufan Bill Li, Mihai Nica, Daniel M. RoyNeurIPS 2022 · 51 citations
- Neural Tangent Kernel Beyond the Infinite-Width Limit: Effects of Depth and InitializationMariia Seleznova, Gitta KutyniokICML 2022 · 34 citations
- Thompson Sampling in Function Spaces via Neural OperatorsRafael Oliveira, Xuesong Wang, Kian Ming A. Chai, Edwin V. BonillaNeurIPS 2025 · 2 citations
- Optimization and Bayes: A Trade-off for Overparameterized Neural NetworksZhengmian Hu, Heng HuangNeurIPS 2023 · 1 citation
Builds on6
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov et al.ICLR 2020 · 167 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 107 citations
- Deep Networks and the Multiple Manifold ProblemSam Buchanan, Dar Gilboa, John WrightICLR 2021 · 9 citations
Related papers
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 101 citations
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 1 citation
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- Divergence of Neural Tangent Kernel in Classification ProblemsZixiong Yu, Songtao Tian, Guhan ChenICLR 2025
- Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural NetworksChaoyue Liu, Han Bi, Like Hui, Xiao LiuNeurIPS 2025
