Scaling Neural Tangent Kernels via Sketching and Random Features
Amir Zandieh, Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, Jinwoo Shin
摘要
The Neural Tangent Kernel (NTK) characterizes the behavior of infinitely-wide neural networks trained under least squares loss by gradient descent. Recent works also report that NTK regression can outperform finitely-wide neural networks trained on small-scale datasets. However, the computational complexity of kernel methods has limited its use in large-scale learning tasks. To accelerate learning with NTK, we design a near input-sparsity time approximation algorithm for NTK, by sketching the polynomial expansions of arc-cosine kernels: our sketch for the convolutional counterpart of NTK (CNTK) can transform any image using a linear runtime in the number of pixels. Furthermore, we prove a spectral approximation guarantee for the NTK matrix, by combining random features (based on leverage score sampling) of the arc-cosine kernels with a sketching algorithm. We benchmark our methods on various large-scale regression and classification tasks and show that a linear regressor trained on our CNTK features matches the accuracy of exact CNTK on CIFAR-10 dataset while achieving 150× speedup. However, the NTK-based approaches encounter the computational bottlenecks of kernel learning. In particular, for a dataset of n images x 1 , x 2 , . . . x n ∈ R d×d , only writing down the CNTK kernel matrix requires Ω d 4 • n 2 operations [5] . Running regression or PCA on the resulting kernel matrix takes additional cubic time in n, which is infeasible in large-scale setups.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 被引用 313 次
- Efficient Dataset Distillation using Random Feature ApproximationNoel Loo, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2022 · 被引用 156 次
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationNanyang Ye, Kaican Li, Haoyue Bai, Runpeng Yu 等CVPR 2022 · 被引用 74 次
- Deep Active Learning by Leveraging Training DynamicsHaonan Wang, Wei Huang, Ziwei Wu, Hanghang Tong 等NeurIPS 2022 · 被引用 49 次
- Fast Neural Kernel Embeddings for General ActivationsInsu Han, Amir Zandieh, Jaehoon Lee, Roman Novak 等NeurIPS 2022 · 被引用 26 次
它引用的顶会 Paper9
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 被引用 147 次
- On the Similarity between the Laplace and Neural Tangent KernelsAmnon Geifman, Abhay Kumar Yadav, Yoni Kasten, Meirav Galun 等NeurIPS 2020 · 被引用 118 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
- Generalized Leverage Score Sampling for Neural NetworksJason D. Lee, Ruoqi Shen, Zhao Song, Mengdi Wang 等NeurIPS 2020 · 被引用 44 次
相关 Paper
- Fast Graph Neural Tangent Kernel via Kronecker SketchingShunhua Jiang, Yunze Man, Zhao Song, Zheng Yu 等AAAI 2022 · 被引用 9 次
- On the Spectral Differences Between NTK and CNTK and Their Implications for Point Cloud RecognitionYuanqu Mou, Chang Gou, Haiyang Bai, Jia LiuICLR 2026
- Fast Sketching of Polynomial Kernels of Polynomial DegreeZhao Song, David P. Woodruff, Zheng Yu, Lichen ZhangICML 2021 · 被引用 48 次
- A Fast, Well-Founded Approximation to the Empirical Neural Tangent KernelMohamad Amin Mohamadi, Wonho Bae, Danica J. SutherlandICML 2023 · 被引用 34 次
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 被引用 72 次
