PseuZO: Pseudo-Zeroth-Order Algorithm for Training Deep Neural Networks
Pengyun Yue, Xuanlin Yang, Mingqing Xiao, Zhouchen Lin
Abstract
Zeroth-order Optimization (ZO) has received wide attention in machine learning, especially when computing full gradient is expensive or even impossible. Recently, ZO has emerged as an important paradigm for memory-efficient fine-tuning of large language models (LLMs), circumventing the memory overhead of backpropagation. However, existing ZO gradient estimators exhibit dimension-dependent variance scaling as !(d), leading to dimension-dependent convergence rates without further assumptions on the objective function, which is prohibitive for large-scale LLM parameters. To address this problem, we present a Pseudo-Zeroth-Order (PseuZO) framework for optimizing composite objective functions, especially large-scale models: min x→X F(x) = E z g → h(x; z), where h represents complex, high-dimensional representations and g is a task-specific loss. While existing zeroth-order methods estimate gradients with final loss functions, our PseuZO algorithm estimate the Jacobian matrix of h(x) with the model output o = h(x), and the gradient of the loss function on model output e = ↑ o g(o), and apply exponential moving average on Jacobian estimators to reduce the variance. Moreover, we use the sliding window technique to reduce memory costs. Our algorithm achieves an O max ω 1 Lε ↑2
, ω 1 Lϑ 2 2 ε ↑4 convergence rate, where ω 1 is the effective dimension of F. Experimental results demonstrate that PseuZO outperforms MeZO and MeZO-SVRG in classification, multiple choice and generation tasks in both full-parameter and PEFT fine-tuning settings by boosting convergence in the early stages of training. For instance, under the same computation time, with respect to SST2 task, PesuZO gets 9.8% higher accuracy than MeZO (91.2% v.s. 82.4%). With the sliding window technique, our PseuZO achieves 70% ↓ 80% memory reduction compared to FO-SGD for different model sizes as PseuZO only introduced a small dimension-independent memory overhead, which enables efficient scaling of the model size. The code is available at https://github.com/YangBigMn/PseuZO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9094b033-35c8-4cd2-888f-36640cde88a1Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
Related papers
- Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank StructuresYiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan et al.ICLR 2025
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman et al.ICML 2024 · 45 citations
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li et al.ICCV 2025 · 3 citations
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
