Lune

ICML2026Top-tier venue

Preconditioning Neural Tangent Kernel for Adaptive Optimization

Xiyuan Yang, Wenxuan Bao, Katherine Tieu, Jingrui He

2026Year

Abstract

The Neural Tangent Kernel is a theoretical framework for understanding the training dynamics of neural networks. However, standard NTK and its variants fail to properly depict the finetuning of foundation models, as they neglect the preconditioning effects of adaptive gradients. To bridge this gap, we propose the Optimizer Aware Kernel (OAK), which incorporates the optimizer's adaptivity into standard NTK framework by a preconditioner estimation technique. Beyond the proposed method, we investigate a fundamental theoretical issue in the field: When and how the kernel regime collapses in finetuning. We derive explicit error bounds showing that the collapse of kernel regime is primarily due to the cumulative training effects and the task discrepancy between pretraining and finetuning. Theoretically, we justify OAK's preconditioner estimation by bounding its error term. Empirically, experiments on various model architectures show both the effectiveness of the OAK method and validity of our arguments on kernel regime collapse. Code is available at https://github.com/xiyuanyang45/ Optimizer-Aware-Kernels.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext becd67a2-3c3d-4868-b16c-79a699e4d2d2

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines