High-Dimensional Distributed Sparse Classification with Scalable Communication-Efficient Global Updates
Fred Lu, Ryan R. Curtin, Edward Raff, Francis Ferraro, James Holt
Abstract
As the size of datasets used in statistical learning continues to grow, distributed training of models has attracted increasing attention. These methods partition the data and exploit parallelism to reduce memory and runtime, but suffer increasingly from communication costs as the data size or the number of iterations grows. Recent work on linear models has shown that a surrogate likelihood can be optimized locally to iteratively improve on an initial solution in a communication-efficient manner. However, existing versions of these methods experience multiple shortcomings as the data size becomes massive, including diverging updates and efficiently handling sparsity. In this work we develop solutions to these problems which enable us to learn a communication-efficient distributed logistic regression model even beyond millions of features. In our experiments we demonstrate a large improvement in accuracy over distributed algorithms with only a few distributed update steps needed, and similar or faster runtimes. Our code is available at https://github.com/FutureComputing4AI/ProxCSL . CCS Concepts • Computing methodologies → Distributed algorithms; Learning linear models; • Mathematics of computing → Multivariate statistics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d67a36c-bc4d-4325-863a-9d674393b120Cited by top-tier papers1
Ask how each one uses itBuilds on3
- A General Framework for Auditing Differentially Private Machine LearningFred Lu, Joseph Munoz, Maya Fuchs, Tyler LeBlond et al.NeurIPS 2022 · 57 citations
- Scaling Up Differentially Private LASSO Regularized Logistic Regression via Faster Frank-Wolfe IterationsEdward Raff, Amol Khanna, Fred LuNeurIPS 2023 · 11 citations
- A Coreset Learning Reality CheckFred Lu, Edward Raff, James HoltAAAI 2023 · 5 citations
Related papers
- Variance Reduced ProxSkip: Algorithm, Theory and Application to Federated LearningGrigory Malinovsky, Kai Yi, Peter RichtárikNeurIPS 2022 · 52 citations
- Almost Linear Constant-Factor Sketching for and Logistic RegressionAlexander Munteanu, Simon Omlor, David P. WoodruffICLR 2023
- Local Steps Speed Up Local GD for Heterogeneous Distributed Logistic RegressionMichael Crawshaw, Blake Woodworth, Mingrui LiuICLR 2025
- FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient EstimationSrijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi et al.ICML 2025
- Distributed Randomized Sketching Kernel LearningRong Yin, Yong Liu, Dan MengAAAI 2022 · 4 citations
