High-Dimensional Distributed Sparse Classification with Scalable Communication-Efficient Global Updates
Fred Lu, Ryan R. Curtin, Edward Raff, Francis Ferraro, James Holt
摘要
As the size of datasets used in statistical learning continues to grow, distributed training of models has attracted increasing attention. These methods partition the data and exploit parallelism to reduce memory and runtime, but suffer increasingly from communication costs as the data size or the number of iterations grows. Recent work on linear models has shown that a surrogate likelihood can be optimized locally to iteratively improve on an initial solution in a communication-efficient manner. However, existing versions of these methods experience multiple shortcomings as the data size becomes massive, including diverging updates and efficiently handling sparsity. In this work we develop solutions to these problems which enable us to learn a communication-efficient distributed logistic regression model even beyond millions of features. In our experiments we demonstrate a large improvement in accuracy over distributed algorithms with only a few distributed update steps needed, and similar or faster runtimes. Our code is available at https://github.com/FutureComputing4AI/ProxCSL . CCS Concepts • Computing methodologies → Distributed algorithms; Learning linear models; • Mathematics of computing → Multivariate statistics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- A General Framework for Auditing Differentially Private Machine LearningFred Lu, Joseph Munoz, Maya Fuchs, Tyler LeBlond 等NeurIPS 2022 · 被引用 57 次
- Scaling Up Differentially Private LASSO Regularized Logistic Regression via Faster Frank-Wolfe IterationsEdward Raff, Amol Khanna, Fred LuNeurIPS 2023 · 被引用 11 次
- A Coreset Learning Reality CheckFred Lu, Edward Raff, James HoltAAAI 2023 · 被引用 5 次
相关 Paper
- Variance Reduced ProxSkip: Algorithm, Theory and Application to Federated LearningGrigory Malinovsky, Kai Yi, Peter RichtárikNeurIPS 2022 · 被引用 52 次
- Almost Linear Constant-Factor Sketching for and Logistic RegressionAlexander Munteanu, Simon Omlor, David P. WoodruffICLR 2023
- Local Steps Speed Up Local GD for Heterogeneous Distributed Logistic RegressionMichael Crawshaw, Blake Woodworth, Mingrui LiuICLR 2025
- FSL-SAGE: Accelerating Federated Split Learning via Smashed Activation Gradient EstimationSrijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi 等ICML 2025
- Distributed Randomized Sketching Kernel LearningRong Yin, Yong Liu, Dan MengAAAI 2022 · 被引用 4 次
