Toward Large Kernel Models
Amirhesam Abedsoltan, Mikhail Belkin, Parthe Pandit
摘要
Recent studies indicate that kernel machines can often perform similarly or better than deep neural networks (DNNs) on small datasets. The interest in kernel machines has been additionally bolstered by the discovery of their equivalence to wide neural networks in certain regimes. However, a key feature of DNNs is their ability to scale the model size and training data size independently, whereas in traditional kernel machines model size is tied to data size. Because of this coupling, scaling kernel machines to large data has been computationally challenging. In this paper, we provide a way forward for constructing large-scale general kernel models, which are a generalization of kernel machines that decouples the model and data, allowing training on large datasets. Specifically, we introduce EigenPro 3.0, an algorithm based on projected dual preconditioned SGD and show scaling to model and data sizes which have not been possible with existing kernel methods. We provide a PyTorch based implementation which can take advantage of multiple GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Estimating Koopman operators with sketching to provably learn large scale dynamical systemsGiacomo Meanti, Antoine Chatalic, Vladimir Kostic, Pietro Novelli 等NeurIPS 2023 · 被引用 22 次
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular dataDaniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan, Mikhail BelkinICLR 2026 · 被引用 18 次
- A theoretical design of concept sets: improving the predictability of concept bottleneck modelsMax Ruiz Luyten, Mihaela van der SchaarNeurIPS 2024 · 被引用 12 次
- Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningAhmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet AggarwalNeurIPS 2023 · 被引用 8 次
- Optimal Kernel Quantile Learning with Random FeaturesCaixing Wang, Xingdong FengICML 2024 · 被引用 3 次
它引用的顶会 Paper9
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya 等NeurIPS 2022 · 被引用 566 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 被引用 138 次
相关 Paper
- Fast Training of Large Kernel Models with Delayed ProjectionsAmirhesam Abedsoltan, Siyuan Ma, Parthe Pandit, Misha BelkinNeurIPS 2025 · 被引用 2 次
- Adaptive kernel predictors from feature-learning infinite limits of neural networksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICML 2025
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 110 次
- Deep Equals Shallow for ReLU Networks in Kernel RegimesAlberto Bietti, Francis R. BachICLR 2021 · 被引用 9 次
