Toward Large Kernel Models
Amirhesam Abedsoltan, Mikhail Belkin, Parthe Pandit
Abstract
Recent studies indicate that kernel machines can often perform similarly or better than deep neural networks (DNNs) on small datasets. The interest in kernel machines has been additionally bolstered by the discovery of their equivalence to wide neural networks in certain regimes. However, a key feature of DNNs is their ability to scale the model size and training data size independently, whereas in traditional kernel machines model size is tied to data size. Because of this coupling, scaling kernel machines to large data has been computationally challenging. In this paper, we provide a way forward for constructing large-scale general kernel models, which are a generalization of kernel machines that decouples the model and data, allowing training on large datasets. Specifically, we introduce EigenPro 3.0, an algorithm based on projected dual preconditioned SGD and show scaling to model and data sizes which have not been possible with existing kernel methods. We provide a PyTorch based implementation which can take advantage of multiple GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 035d7ea1-7e0e-4119-a22d-904b00d47b4eCited by top-tier papers8
- Estimating Koopman operators with sketching to provably learn large scale dynamical systemsGiacomo Meanti, Antoine Chatalic, Vladimir Kostic, Pietro Novelli et al.NeurIPS 2023 · 22 citations
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular dataDaniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan, Mikhail BelkinICLR 2026 · 18 citations
- A theoretical design of concept sets: improving the predictability of concept bottleneck modelsMax Ruiz Luyten, Mihaela van der SchaarNeurIPS 2024 · 12 citations
- Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningAhmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet AggarwalNeurIPS 2023 · 8 citations
- Optimal Kernel Quantile Learning with Random FeaturesCaixing Wang, Xingdong FengICML 2024 · 3 citations
Builds on9
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov et al.ICLR 2020 · 167 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
Related papers
- Fast Training of Large Kernel Models with Delayed ProjectionsAmirhesam Abedsoltan, Siyuan Ma, Parthe Pandit, Misha BelkinNeurIPS 2025 · 2 citations
- Adaptive kernel predictors from feature-learning infinite limits of neural networksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICML 2025
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 110 citations
- Deep Equals Shallow for ReLU Networks in Kernel RegimesAlberto Bietti, Francis R. BachICLR 2021 · 9 citations
