Fast Training of Large Kernel Models with Delayed Projections
Amirhesam Abedsoltan, Siyuan Ma, Parthe Pandit, Misha Belkin
Abstract
Classical kernel machines have historically faced significant challenges in scaling to large datasets and model sizes--a key ingredient that has driven the success of neural networks. In this paper, we present a new methodology for building kernel machines that can scale efficiently with both data size and model size. Our algorithm introduces delayed projections to Preconditioned Stochastic Gradient Descent (PSGD) allowing the training of much larger models than was previously feasible, pushing the practical limits of kernel-based learning. We validate our algorithm, EigenPro4, across multiple datasets, demonstrating drastic training speed up over the existing methods while maintaining comparable or better classification accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 199 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- The Deep Bootstrap Framework: Good Online Learners are Good Offline GeneralizersPreetum Nakkiran, Behnam Neyshabur, Hanie SedghiICLR 2021 · 75 citations
- Toward Large Kernel ModelsAmirhesam Abedsoltan, Mikhail Belkin, Parthe PanditICML 2023 · 23 citations
Related papers
- Nonlinear Sufficient Dimension Reduction with a Stochastic Neural NetworkSiqi Liang, Yan Sun, Faming LiangNeurIPS 2022 · 20 citations
- Preconditioning for Scalable Gaussian Process Hyperparameter OptimizationJonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham et al.ICML 2022 · 36 citations
- Fast and Scalable Adversarial Training of Kernel SVM via Doubly Stochastic GradientsHuimin Wu, Zhengmian Hu, Bin GuAAAI 2021 · 10 citations
- Improving Neural Network Training in Low Dimensional Random BasesFrithjof Gressmann, Zach Eaton-Rosen, Carlo LuschiNeurIPS 2020 · 35 citations
- Learning with Optimized Random Features: Exponential Speedup by Quantum Machine Learning without Sparsity and Low-Rank AssumptionsHayata Yamasaki, Sathyawageeswar Subramanian, Sho Sonoda, Masato KoashiNeurIPS 2020 · 23 citations
