Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks
Russell Tsuchida, Tim Pearce, Christopher van der Heide, Fred Roosta, Marcus Gallagher
摘要
Analysing and computing with Gaussian processes arising from infinitely wide neural networks has recently seen a resurgence in popularity. Despite this, many explicit covariance functions of networks with activation functions used in modern networks remain unknown. Furthermore, while the kernels of deep networks can be computed iteratively, theoretical understanding of deep kernels is lacking, particularly with respect to fixed-point dynamics. Firstly, we derive the covariance functions of multi-layer perceptrons (MLPs) with exponential linear units (ELU) and Gaussian error linear units (GELU) and evaluate the performance of the limiting Gaussian processes on some benchmarks. Secondly, and more generally, we analyse the fixed-point dynamics of iterated kernels corresponding to a broad range of activation functions. We find that unlike some previously studied neural network kernels, these new kernels exhibit non-trivial fixed-point dynamics which are mirrored in finite-width neural networks. The fixed point behaviour present in some networks explains a mechanism for implicit regularisation in overparameterised deep models. Our results relate to both the static iid parameter conjugate kernel and the dynamic neural tangent kernel constructions 1 . where n is the number of neurons in the hidden layer and the output bias satisfies V b ∼ N (0, σ 2 b ). The output evaluated at input x 1 is
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fast Neural Kernel Embeddings for General ActivationsInsu Han, Amir Zandieh, Jaehoon Lee, Roman Novak 等NeurIPS 2022 · 被引用 26 次
- Squared Neural Families: A New Class of Tractable Density ModelsRussell Tsuchida, Cheng Soon Ong, Dino SejdinovicNeurIPS 2023 · 被引用 15 次
- Critical Initialization of Wide and Deep Neural Networks using Partial Jacobians: General Theory and ApplicationsDarshil Doshi, Tianyu He, Andrey GromovNeurIPS 2023 · 被引用 10 次
- Deep Kernel Posterior Learning under Infinite Variance Prior WeightsJorge Loría, Anindya BhadraICLR 2025
它引用的顶会 Paper7
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 被引用 107 次
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 被引用 81 次
相关 Paper
- An Infinite-Width Analysis on the Jacobian-Regularised Training of a Neural NetworkTaeyoung Kim, Hongseok YangICML 2024 · 被引用 2 次
- On the Spectral Bias of Convolutional Neural Tangent and Gaussian Process KernelsAmnon Geifman, Meirav Galun, David Jacobs, Ronen BasriNeurIPS 2022 · 被引用 22 次
- Stationary Activations for Uncertainty Calibration in Deep LearningLassi Meronen, Christabella Irwanto, Arno SolinNeurIPS 2020 · 被引用 22 次
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 被引用 1 次
- Deep Equals Shallow for ReLU Networks in Kernel RegimesAlberto Bietti, Francis R. BachICLR 2021 · 被引用 9 次
