GPU-Accelerated Primal Learning for Extremely Fast Large-Scale Classification
John T. Halloran, David M. Rocke
摘要
One of the most efficient methods to solve L 2 -regularized primal problems, such as logistic regression and linear support vector machine (SVM) classification, is the widely used trust region Newton algorithm, TRON [38]. While TRON has recently been shown to enjoy substantial speedups on shared-memory multi-core systems [35, 22] , exploiting graphical processing units (GPUs) to speed up the method is significantly more difficult, owing to the highly complex and heavily sequential nature of the algorithm. In this work, we show that using judicious GPU-optimization principles, TRON training time for different losses and feature representations may be drastically reduced. For sparse feature sets, we show that using GPUs to train logistic regression classifiers in LIBLINEAR is up to an orderof-magnitude faster than solely using multithreading. For dense feature sets-which impose far more stringent memory constraints-we show that GPUs substantially reduce the lengthy SVM learning times required for state-of-the-art proteomics analysis, leading to dramatic improvements over recently proposed speedups. Furthermore, we show how GPU speedups may be mixed with multithreading to enable such speedups when the dataset is too large for GPU memory requirements; on a massive dense proteomics dataset of nearly a quarter-billion data instances, these mixed-architecture speedups reduce SVM analysis time from over half a week to less than a single day while using limited GPU memory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Differentiable Optimization of Generalized Nondecomposable Functions using Linear ProgramsZihang Meng, Lopamudra Mukherjee, Yichao Wu, Vikas Singh 等NeurIPS 2021 · 被引用 1 次
- Giga-scale Kernel Matrix-Vector Multiplication on GPURobert Hu, Siu Lun Chau, Dino Sejdinovic, Joan GlaunèsNeurIPS 2022 · 被引用 3 次
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- Beyond L1: Faster and Better Sparse Models with skglmQuentin Bertrand, Quentin Klopfenstein, Pierre-Antoine Bannier, Gauthier Gidel 等NeurIPS 2022 · 被引用 32 次
- PRISM: Distribution-free Adaptive Computation of Matrix Functions for Accelerating Neural Network TrainingShenghao Yang, Zhichao Wang, Oleg Balabanov, N. Benjamin Erichson 等ICML 2026 · 被引用 3 次
