ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order Optimization
Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang, Yukang Lin, Xiangping Wu, Chuanyi Liu, Xiaobao Song
摘要
Lowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng 等NeurIPS 2025 · 被引用 66 次
- FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale SpeedSizhe Dang, yangyangGuo, Yanjun Zhao, Xiaodong Zheng 等ICLR 2026 · 被引用 16 次
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等NeurIPS 2025 · 被引用 8 次
- AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-TuningYifan Yang, Kai Zhen, Ershad Banijamali, Athanasios Mouchtaris 等EMNLP 2024 · 被引用 7 次
- Towards Efficient Low-Order Hybrid Optimizer for Language Model Fine-TuningMinping Chen, You-Liang Huang, Zeyi WenAAAI 2025 · 被引用 6 次
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda 等NeurIPS 2020 · 被引用 697 次
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam 等AAAI 2023 · 被引用 234 次
相关 Paper
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li 等ICCV 2025 · 被引用 3 次
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the MomentsZhijie Cai, Haolong Chen, Guangxu ZhuICML 2026
- Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language ModelsZeman Li, Xinwei Zhang, Peilin Zhong, Yuan Deng 等ICLR 2025
- Second-Order Fine-Tuning without Pain for LLMs: A Hessian Informed Zeroth-Order OptimizerYanjun Zhao, Sizhe Dang, Haishan Ye, Guang Dai 等ICLR 2025
