Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
Shayan Talaei, Matin Ansaripour, Giorgi Nadiradze, Dan Alistarh
摘要
Distributed optimization is the standard way of speeding up machine learning training, and most of the research in the area focuses on distributed first-order, gradient-based methods. Yet, there are settings where some computationally-bounded nodes may not be able to implement first-order, gradient-based optimization, while they could still contribute to joint optimization tasks. In this paper, we initiate the study of hybrid decentralized optimization, studying settings where nodes with zeroth-order and first-order optimization capabilities co-exist in a distributed system, and attempt to jointly solve an optimization task over some data distribution. We essentially show that, under reasonable parameter settings, such a system can not only withstand noisier zeroth-order agents but can even benefit from integrating such agents into the optimization process, rather than ignoring their information. At the core of our approach is a new analysis of distributed optimization with noisy and possibly-biased gradient estimators, which may be of independent interest. Our results hold for both convex and non-convex objectives. Experimental results on standard optimization tasks confirm our analysis, showing that hybrid first-zeroth order optimization can be practical, even when training deep neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language ModelsZiwei Liu, Borui Kang, Wei Li, Hangjie Yuan 等AAAI 2026
- HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free AggregationQiyuan Chen, Xian Wu, Yi Wang, Xianhao ChenICML 2026
它引用的顶会 Paper3
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Decentralized Deep Learning with Arbitrary Communication CompressionAnastasia Koloskova, Tao Lin, Sebastian U. Stich, Martin JaggiICLR 2020 · 被引用 263 次
- Single Point-Based Distributed Zeroth-Order Optimization with a Non-Convex Stochastic Objective FunctionElissa Mhanna, Mohamad AssaadICML 2023 · 被引用 10 次
相关 Paper
- A Zeroth-Order ADMM Algorithm for Stochastic Optimization over Distributed Processing NetworksZai Shi, Atilla EryilmazINFOCOM 2020 · 被引用 4 次
- A Hybrid Variance-Reduced Method for Decentralized Stochastic Non-Convex OptimizationRan Xin, Usman A. Khan, Soummya KarICML 2021 · 被引用 51 次
- Distributed Zero-Order Optimization under Adversarial NoiseArya Akhavan, Massimiliano Pontil, Alexandre B. TsybakovNeurIPS 2021 · 被引用 28 次
- Decentralized Sporadic Federated Learning: A Unified Algorithmic Framework with Convergence GuaranteesShahryar Zehtabi, Dong-Jun Han, Rohit Parasnis, Seyyedali Hosseinalipour 等ICLR 2025
- Adaptive Random Walk Gradient Descent for Decentralized OptimizationTao Sun, Dongsheng Li, Bao WangICML 2022 · 被引用 24 次
