Refining Adaptive Zeroth-Order Optimization at Ease
Yao Shu, Qixin Zhang, Kun He, Zhongxiang Dai
Abstract
Recently, zeroth-order (ZO) optimization plays an essential role in scenarios where gradient information is inaccessible or unaffordable, such as blackbox systems and resource-constrained environments. While existing adaptive methods such as ZO-AdaMM have shown promise, they are fundamentally limited by their underutilization of moment information during optimization, usually resulting in underperforming convergence. To overcome these limitations, this paper introduces Refined Adaptive Zeroth-Order Optimization (R-AdaZO). Specifically, we first show the untapped variance reduction effect of first moment estimate on ZO gradient estimation, which improves the accuracy and stability of ZO updates. We then refine the second moment estimate based on these variance-reduced gradient estimates to better capture the geometry of the optimization landscape, enabling a more effective scaling of ZO updates. We present rigorous theoretical analysis to show (I) the first analysis to the variance reduction of first moment estimate in ZO optimization, (II) the improved second moment estimates with a more accurate approximation of its variance-free ideal, (III) the first variance-aware convergence framework for adaptive ZO optimizers, which may be of independent interest, and (IV) the faster convergence of R-AdaZO than existing baselines like ZO-AdaMM. Our extensive experiments, including synthetic problems, black-box adversarial attack, and memory-efficient fine-tuning of large language models (LLMs), further verify the superior convergence of R-AdaZO, indicating that R-AdaZO offers an improved solution for realworld ZO optimization challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6eab0e6-3ffa-43dc-b281-381cf925e690Cited by top-tier papers3
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM UnlearningYicheng Lang, Yihua Zhang, Chongyu Fan, Changsheng Wang et al.ICLR 2026 · 4 citations
- Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization LensJunbin Qiu, Zhaowei Hong, Renzhe Xu, Yao ShuICML 2026
- PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient TrainingYanyi Li, Yimu Zhang, Cong FangICML 2026
Builds on8
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li et al.ICML 2024 · 134 citations
- BayesOpt Adversarial AttackBinxin Ru, Adam D. Cobb, Arno Blaas, Yarin GalICLR 2020 · 85 citations
- On the Convergence of Prior-Guided Zeroth-Order Optimization AlgorithmsShuyu Cheng, Guoqiang Wu, Jun ZhuNeurIPS 2021 · 27 citations
Related papers
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the MomentsZhijie Cai, Haolong Chen, Guangxu ZhuICML 2026
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li et al.ICCV 2025 · 3 citations
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman et al.ICML 2024 · 45 citations
- ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order OptimizationShuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang et al.AAAI 2024 · 27 citations
