Preconditioning Neural Tangent Kernel for Adaptive Optimization
Xiyuan Yang, Wenxuan Bao, Katherine Tieu, Jingrui He
Abstract
The Neural Tangent Kernel is a theoretical framework for understanding the training dynamics of neural networks. However, standard NTK and its variants fail to properly depict the finetuning of foundation models, as they neglect the preconditioning effects of adaptive gradients. To bridge this gap, we propose the Optimizer Aware Kernel (OAK), which incorporates the optimizer's adaptivity into standard NTK framework by a preconditioner estimation technique. Beyond the proposed method, we investigate a fundamental theoretical issue in the field: When and how the kernel regime collapses in finetuning. We derive explicit error bounds showing that the collapse of kernel regime is primarily due to the cumulative training effects and the task discrepancy between pretraining and finetuning. Theoretically, we justify OAK's preconditioner estimation by bounding its error term. Empirically, experiments on various model architectures show both the effectiveness of the OAK method and validity of our arguments on kernel regime collapse. Code is available at https://github.com/xiyuanyang45/ Optimizer-Aware-Kernels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext becd67a2-3c3d-4868-b16c-79a699e4d2d2Builds on10
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov et al.ICLR 2020 · 167 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 90 citations
Related papers
- Label-Aware Neural Tangent Kernel: Toward Better Generalization and Local ElasticityShuxiao Chen, Hangfeng He, Weijie J. SuNeurIPS 2020 · 25 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Neural (Tangent Kernel) CollapseMariia Seleznova, Dana Weitzner, Raja Giryes, Gitta Kutyniok et al.NeurIPS 2023 · 23 citations
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- Neural Tangent Kernels Under Stochastic Data AugmentationJoshua DeOliveira, Sajal Chakroborty, Walter Gerych, Elke A. RundensteinerAAAI 2026
