Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks
Chaoyue Liu, Han Bi, Like Hui, Xiao Liu
摘要
Nonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this work, we focus on ReLU activation and reveal a novel and intriguing property of nonlinear activations. By comparing enabling and disabling the nonlinear activations in the neural network, we demonstrate their specific effects on wide neural networks: (a) better feature separation, i.e., a larger angle separation for similar data in the feature space of model gradient, and (b) better NTK conditioning, i.e., a smaller condition number of neural tangent kernel (NTK). Furthermore, we show that the network depth (i.e., with more nonlinear activation operations) further amplifies these effects; in addition, in the infinite-width-then-depth limit, all data are equally separated with a fixed angle in the model gradient feature space, regardless of how similar they are originally in the input space. Note that, without the nonlinear activation, i.e., in a linear neural network, the data separation remains the same as for the original inputs and NTK condition number is equivalent to the Gram matrix, regardless of the network depth. Due to the close connection between NTK condition number and convergence theories, our results imply that nonlinear activation helps to improve the worst-case convergence rates of gradient based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun 等NeurIPS 2020 · 被引用 118 次
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 被引用 107 次
- Spectra of the Conjugate Kernel and Neural Tangent Kernel for linear-width neural networksZhou Fan, Zhichao WangNeurIPS 2020 · 被引用 101 次
相关 Paper
- Fast Neural Kernel Embeddings for General ActivationsInsu Han, Amir Zandieh, Jaehoon Lee, Roman Novak 等NeurIPS 2022 · 被引用 26 次
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 被引用 1 次
- On the Random Conjugate Kernel and Neural Tangent KernelZhengmian Hu, Heng HuangICML 2021 · 被引用 15 次
- Characterizing the spectrum of the NTK via a power series expansionMichael Murray, Hui Jin, Benjamin Bowman, Guido MontúfarICLR 2023 · 被引用 2 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
