A Fast, Well-Founded Approximation to the Empirical Neural Tangent Kernel
Mohamad Amin Mohamadi, Wonho Bae, Danica J. Sutherland
Abstract
Empirical neural tangent kernels (eNTKs) can provide a good understanding of a given network's representation: they are often far less expensive to compute and applicable more broadly than infinite width NTKs. For networks with O output units (e.g. an O-class classifier), however, the eNTK on N inputs is of size , taking memory and up to computation. Most existing applications have therefore used one of a handful of approximations yielding kernel matrices, saving orders of magnitude of computation, but with limited to no justification. We prove that one such approximation, which we call"sum of logits", converges to the true eNTK at initialization for any network with a wide final"readout"layer. Our experiments demonstrate the quality of this approximation for various uses across a range of settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad766dc0-1e0c-41bb-96eb-7f38f29e54b5Cited by top-tier papers18
- Understanding Linear Probing then Fine-tuning Language Models from NTK PerspectiveAkiyoshi Tomihari, Issei SatoNeurIPS 2024 · 30 citations
- Heterogeneous Personalized Federated Learning by Local-Global Updates Mixing via Convergence RateMeirui Jiang, Anjie Le, Xiaoxiao Li, Qi DouICLR 2024 · 13 citations
- Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondAlan Jeffares, Alicia Curth, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate ModelsAndrew Engel, Zhichao Wang, Natalie Frank, Ioana Dumitriu et al.ICLR 2024 · 8 citations
- lpNTK: Better Generalisation with Less Data via Sample Interaction During LearningShangmin Guo, Yi Ren, Stefano V. Albrecht, Kenny SmithICLR 2024 · 7 citations
Builds on18
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
- Dataset Meta-Learning from Kernel Ridge-RegressionTimothy Nguyen, Zhourong Chen, Jaehoon LeeICLR 2021 · 307 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
Related papers
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 72 citations
- The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich RegimesAlexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz PehlevanICLR 2023 · 4 citations
- Training-Free Determination of Network Width via Neural Tangent KernelTatsumi Sunada, Toshihiko Yamasaki, Atsuto MakiICLR 2026
- Determinant Estimation under Memory Constraints and Neural Scaling LawsSiavash Ameli, Chris van der Heide, Liam Hodgkinson, Fred Roosta et al.ICML 2025
