Statistically Preconditioned Accelerated Gradient Method for Distributed Optimization
Hadrien Hendrikx, Lin Xiao, Sébastien Bubeck, Francis R. Bach, Laurent Massoulié
摘要
We consider the setting of distributed empirical risk minimization where multiple machines compute the gradients in parallel and a centralized server updates the model parameters. In order to reduce the number of communications required to reach a given accuracy, we propose a preconditioned accelerated gradient method where the preconditioning is done by solving a local optimization problem over a subsampled dataset at the server. The convergence rate of the method depends on the square root of the relative condition number between the global and local loss functions. We estimate the relative condition number for linear prediction models by studying uniform concentration of the Hessians over a bounded domain, which allows us to derive improved convergence rates for existing preconditioned gradient methods and our accelerated method. Experiments on real-world datasets illustrate the benefits of acceleration in the ill-conditioned regime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Decentralized Local Stochastic Extra-Gradient for Variational InequalitiesAleksandr Beznosikov, Pavel E. Dvurechensky, Anastasia Koloskova, Valentin Samokhin 等NeurIPS 2022 · 被引用 49 次
- Distributed Saddle-Point Problems Under Data SimilarityAleksandr Beznosikov, Gesualdo Scutari, Alexander Rogozin, Alexander V. GasnikovNeurIPS 2021 · 被引用 39 次
- Optimal Gradient Sliding and its Application to Optimal Distributed Optimization Under SimilarityDmitry Kovalev, Aleksandr Beznosikov, Ekaterina Borodich, Alexander V. Gasnikov 等NeurIPS 2022 · 被引用 26 次
- Newton Method over Networks is Fast up to the Statistical PrecisionAmir Daneshmand, Gesualdo Scutari, Pavel E. Dvurechensky, Alexander V. GasnikovICML 2021 · 被引用 22 次
- Federated Optimization with Doubly Regularized Drift CorrectionXiaowen Jiang, Anton Rodomanov, Sebastian U. StichICML 2024 · 被引用 18 次
相关 Paper
- Communication-Efficient Distributed Optimization with Quantized PreconditionersFoivos Alimisis, Peter Davies, Dan AlistarhICML 2021 · 被引用 17 次
- Accelerating SGD for Highly Ill-Conditioned Huge-Scale Online Matrix CompletionJialun Zhang, Hong-Ming Chiu, Richard Y. ZhangNeurIPS 2022 · 被引用 12 次
- Randomized Block-Diagonal Preconditioning for Parallel LearningCelestine Mendler-Dünner, Aurélien LucchiICML 2020 · 被引用 1 次
- Do Subsampled Newton Methods Work for High-Dimensional Data?Xiang Li, Shusen Wang, Zhihua ZhangAAAI 2020 · 被引用 15 次
- Communication-Efficient Distributed PCA by Riemannian OptimizationLong-Kai Huang, Sinno Jialin PanICML 2020 · 被引用 22 次
