Distributed Optimization for Overparameterized Problems: Achieving Optimal Dimension Independent Communication Complexity
Bingqing Song, Ioannis C. Tsaknakis, Chung-Yiu Yau, Hoi-To Wai, Mingyi Hong
摘要
Decentralized optimization are playing an important role in applications such as training large machine learning models, among others. Despite its superior practical performance, there has been some lack of fundamental understanding about its theoretical properties. In this work, we address the following open research question: To train an overparameterized model over a set of distributed nodes, what is the minimum communication overhead (in terms of the bits got exchanged) that the system needs to sustain, while still achieving (near) zero training loss? We show that for a class of overparameterized models where the number of parameters D is much larger than the total data samples N , the best possible communication complexity is ⌦(N ), which is independent of the problem dimension D. Further, for a few specific overparameterized models (i.e., the linear regression, and certain multi-layer neural network with one wide layer), we develop a set of algorithms which uses certain linear compression followed by adaptive quantization, and show that they achieve dimension independent, near-optimal communication complexity. To our knowledge, this is the first time that dimension independent communication complexity has been shown for distributed optimization.
Recently, overparameterized problem has drawn great research interests. It has been showed that remarkable performance can be achieved with overparameterized networks [
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?Yutong He, Xinmeng Huang, Kun YuanNeurIPS 2023 · 被引用 25 次
- Understanding Convergence and Generalization in Federated Learning through Feature Learning TheoryWei Huang, Ye Shi, Zhongyi Cai, Taiji SuzukiICLR 2024 · 被引用 17 次
它引用的顶会 Paper11
- EF21: A New, Simpler, Theoretically Better, and Practically Faster Error FeedbackPeter Richtárik, Igor Sokolov, Ilyas FatkhullinNeurIPS 2021 · 被引用 219 次
- MARINA: Faster Non-Convex Distributed Learning with CompressionEduard Gorbunov, Konstantin Burlachenko, Zhize Li, Peter RichtárikICML 2021 · 被引用 129 次
- Optimal Complexity in Decentralized TrainingYucheng Lu, Christopher De SaICML 2021 · 被引用 95 次
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 被引用 82 次
- FL-NTK: A Neural Tangent Kernel-based Framework for Federated Learning AnalysisBaihe Huang, Xiaoxiao Li, Zhao Song, Xin YangICML 2021 · 被引用 66 次
相关 Paper
- Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order OptimizationZhe Li, Bicheng Ying, Zidong Liu, Chaosheng Dong 等ICLR 2025
- Practical Low-Rank Communication Compression in Decentralized Deep LearningThijs Vogels, Sai Praneeth Karimireddy, Martin JaggiNeurIPS 2020 · 被引用 63 次
- Towards Tight Communication Lower Bounds for Distributed OptimisationJanne H. Korhonen, Dan AlistarhNeurIPS 2021 · 被引用 10 次
- Quantized Decentralized Stochastic Learning over Directed GraphsHossein Taheri, Aryan Mokhtari, Hamed Hassani, Ramtin PedarsaniICML 2020 · 被引用 59 次
- Theoretically Better and Numerically Faster Distributed Optimization with Smoothness-Aware Quantization TechniquesBokun Wang, Mher Safaryan, Peter RichtárikNeurIPS 2022 · 被引用 13 次
