Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
Zeman Li, Xinwei Zhang, Peilin Zhong, Yuan Deng, Meisam Razaviyayn, Vahab Mirrokni
Abstract
Fine-tuning language models (LMs) with the standard Adam optimizer often demands excessive memory, limiting accessibility. The "in-place" version of Stochastic Gradient Descent (IP-SGD) and Memory-Efficient Zeroth-order Optimizer (MeZO) have been proposed as solutions to improve memory efficiency. However, IP-SGD still requires a substantial amount of memory, and MeZO suffers from slow convergence and degraded final performance due to its zeroth-order nature. This paper introduces Addax, a novel method that improves both memory efficiency and algorithm performance of IP-SGD by integrating it with MeZO. Specifically, Addax computes the zerothor first-order gradient of the data points in the minibatch based on their memory consumption and combines zeroth-and first-order gradient estimates to obtain the updated direction in each step. By computing the zeroth-order gradient of data points that require more memory and the first-order gradient of the ones that require less memory, Addax overcomes the slow convergence of MeZO and the excessive memory requirement of IP-SGD. Additionally, the zeroth-order gradient acts as a regularizer for the first-order gradient, further enhancing the model's final performance. Theoretically, we establish the convergence of Addax under mild assumptions, demonstrating faster convergence and less restrictive hyper-parameter choices than MeZO. Our extensive experiments with diverse LMs and tasks show that Addax consistently outperforms MeZO in terms of accuracy and convergence speed while having a comparable memory footprint. In particular, our experiments using one A100 GPU on the OPT-13B model reveal that, on average, Addax outperforms MeZO in terms of accuracy/F1 score by 14% and runs 15× faster while having a comparable memory footprint to MeZO. In our experiments on the larger OPT-30B model, on average, Addax outperforms MeZO in terms of accuracy/F1 score by > 16% and runs 30× faster on a single H100 GPU. Moreover, Addax surpasses the performance of standard fine-tuning approaches, such as IP-SGD and Adam, in most tasks in terms of accuracy/F1 score with significantly less memory requirement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh et al.NeurIPS 2025 · 8 citations
- On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order OptimizationShaocong Ma, Heng HuangNeurIPS 2025 · 4 citations
- Private Zeroth-Order Optimization with Public DataXuchen Gong, Tian LiNeurIPS 2025 · 2 citations
- MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language ModelsZhen Zhang, Yifan Yang, Kai Zhen, Nathan Susanj et al.EMNLP 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
Related papers
- Towards Efficient Low-Order Hybrid Optimizer for Language Model Fine-TuningMinping Chen, You-Liang Huang, Zeyi WenAAAI 2025 · 6 citations
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman et al.ICML 2024 · 45 citations
- AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the MomentsZhijie Cai, Haolong Chen, Guangxu ZhuICML 2026
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng et al.NeurIPS 2025 · 66 citations
