Convergence Analysis of the Lion Optimizer in Centralized and Distributed Settings
Wei Jiang, Mao Xu, Wenhao Yang, Yibo Wang, Zechao Li, Lijun Zhang
Abstract
In this paper, we provide a comprehensive convergence analysis for the Lion optimizer. First, we establish that the original Lion achieves a convergence rate of , where denotes the problem dimension and is the iteration number. To improve this rate, we propose a variance reduction variant of Lion, which attains an enhanced rate of with the average smoothness assumption. Then, we extend our analysis to distributed settings. We demonstrate that the distributed Lion optimizer and its variance reduction counterpart achieve linear speedup with respect to the number of nodes , yielding convergence rates of and , respectively. Additionally, we investigate a communication-efficient distributed Lion variant that utilizes sign compression for bidirectional communication. By employing unbiased sign operations, this variant achieves a convergence rate of , and its variance-reduced counterpart can further improve the rate to . Finally, we conduct numerical experiments to validate the effectiveness of the proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de835954-2d62-4cb7-9e37-1fae7e072a04Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex OptimizationZhize Li, Hongyan Bao, Xiangliang Zhang, Peter RichtárikICML 2021 · 164 citations
- STORM+: Fully Adaptive SGD with Recursive Momentum for Nonconvex OptimizationKfir Y. Levy, Ali Kavis, Volkan CevherNeurIPS 2021 · 59 citations
- Momentum Ensures Convergence of SIGNSGD under Weaker AssumptionsTao Sun, Qingsong Wang, Dongsheng Li, Bao WangICML 2023 · 36 citations
Related papers
- Communication Efficient Distributed Training with Distributed LionBo Liu, Lemeng Wu, Lizhang Chen, Kaizhao Liang et al.NeurIPS 2024 · 21 citations
- Efficient Sign-Based Optimization: Accelerating Convergence via Variance ReductionWei Jiang, Sifan Yang, Wenhao Yang, Lijun ZhangNeurIPS 2024 · 19 citations
- Proving the Limited Scalability of Centralized Distributed Optimization via a New Lower Bound ConstructionAlexander TyurinICLR 2026
- Communication Efficient Distributed Newton Method with Fast Convergence RatesChengchang Liu, Lesi Chen, Luo Luo, John C. S. LuiKDD 2023 · 4 citations
- Accelerated Methods with Compressed Communications for Distributed Optimization Problems Under Data SimilarityDmitry Bylinkin, Aleksandr BeznosikovAAAI 2025 · 3 citations
