Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models
Tanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman, Wooseok Ha
摘要
Fine-tuning language models (LMs) has demonstrated success in a wide array of downstream tasks. However, as LMs are scaled up, the memory requirements for backpropagation become prohibitively high. Zeroth-order (ZO) optimization methods can leverage memory-efficient forward passes to estimate gradients. More recently, MeZO, an adaptation of ZO-SGD, has been shown to consistently outperform zero-shot and in-context learning when combined with suitable task prompts. In this work, we couple ZO methods with variance reduction techniques to enhance stability and convergence for inference-based LM fine-tuning. We introduce Memory-Efficient Zeroth-Order Stochastic Variance-Reduced Gradient (MeZO-SVRG) and demonstrate its efficacy across multiple LM fine-tuning tasks, eliminating the reliance on task-specific prompts. Evaluated across a range of both masked and autoregressive LMs on benchmark GLUE tasks, MeZO-SVRG outperforms MeZO with up to 20% increase in test accuracies in both full- and partial-parameter fine-tuning settings. MeZO-SVRG benefits from reduced computation time as it often surpasses MeZO's peak test accuracy with a reduction in GPU-hours. MeZO-SVRG significantly reduces the required memory footprint compared to first-order SGD, i.e. by for autoregressive models. Our experiments highlight that MeZO-SVRG's memory savings progressively improve compared to SGD with larger batch sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuningQitao Tan, Jun Liu, Zheng Zhan, Caiwen Ding 等NeurIPS 2025 · 被引用 19 次
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等NeurIPS 2025 · 被引用 8 次
- AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-TuningYifan Yang, Kai Zhen, Ershad Banijamali, Athanasios Mouchtaris 等EMNLP 2024 · 被引用 7 次
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank StructureBoao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng 等ICLR 2026 · 被引用 7 次
- Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-TrainingReza Shirkavand, Peiran Yu, Qi He, Heng HuangNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang 等ICML 2022 · 被引用 343 次
相关 Paper
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li 等ICCV 2025 · 被引用 3 次
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-TuningYong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng 等NeurIPS 2025 · 被引用 66 次
- PseuZO: Pseudo-Zeroth-Order Algorithm for Training Deep Neural NetworksPengyun Yue, Xuanlin Yang, Mingqing Xiao, Zhouchen LinNeurIPS 2025 · 被引用 5 次
