Fast Training of Large Kernel Models with Delayed Projections
Amirhesam Abedsoltan, Siyuan Ma, Parthe Pandit, Misha Belkin
摘要
Classical kernel machines have historically faced significant challenges in scaling to large datasets and model sizes--a key ingredient that has driven the success of neural networks. In this paper, we present a new methodology for building kernel machines that can scale efficiently with both data size and model size. Our algorithm introduces delayed projections to Preconditioned Stochastic Gradient Descent (PSGD) allowing the training of much larger models than was previously feasible, pushing the practical limits of kernel-based learning. We validate our algorithm, EigenPro4, across multiple datasets, demonstrating drastic training speed up over the existing methods while maintaining comparable or better classification accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 被引用 138 次
- The Deep Bootstrap Framework: Good Online Learners are Good Offline GeneralizersPreetum Nakkiran, Behnam Neyshabur, Hanie SedghiICLR 2021 · 被引用 75 次
- Toward Large Kernel ModelsAmirhesam Abedsoltan, Mikhail Belkin, Parthe PanditICML 2023 · 被引用 23 次
相关 Paper
- Nonlinear Sufficient Dimension Reduction with a Stochastic Neural NetworkSiqi Liang, Yan Sun, Faming LiangNeurIPS 2022 · 被引用 20 次
- Preconditioning for Scalable Gaussian Process Hyperparameter OptimizationJonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham 等ICML 2022 · 被引用 36 次
- Fast and Scalable Adversarial Training of Kernel SVM via Doubly Stochastic GradientsHuimin Wu, Zhengmian Hu, Bin GuAAAI 2021 · 被引用 10 次
- Improving Neural Network Training in Low Dimensional Random BasesFrithjof Gressmann, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2020 · 被引用 35 次
- Learning with Optimized Random Features: Exponential Speedup by Quantum Machine Learning without Sparsity and Low-Rank AssumptionsHayata Yamasaki, Sathyawageeswar Subramanian, Sho Sonoda, Masato KoashiNeurIPS 2020 · 被引用 23 次
