Memory Savings at What Cost? A Study of Alternatives to Backpropagation
Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan
Abstract
Forward-mode automatic differentiation (FMAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagationfree alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing. We present a unified theoretical and empirical comparison of BP, checkpointed BP, FMAD, and ZO for LLM and vision-language model training, showing that while FMAD and ZO reduce activation memory, they trade memory for higher computational cost and longer wallclock time to convergence, resulting in lower accuracy and slower training, especially under constrained perturbation budgets. Across models, BP with checkpointing outperforms FMAD and ZO variants, including variance-reduced methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence, and 3.8× fewer computations at comparable memory usage, while also revealing instability-related failure modes in FMAD and ZO. Overall, our results correct a one-sided benchmarking narrative by showing that memoryefficient methods entail fundamentally different trade-offs, and that ignoring these distinctions has led to misleading conclusions about LLM optimization in prior work. Our source code is available at https://github.com/Astuary/ Gradient_Estimation_Methods . This paper addresses the above-mentioned limitations through a comprehensive study of BP, FMAD, and ZO approaches in the context of LLM training and fine-tuning. We first outline the expected trade-offs among convergence behavior, memory consumption, and computational cost as functions of model dimensionality d and the number
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4260739-0083-457c-a36b-926ea4390e4bBuilds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li et al.ICML 2024 · 134 citations
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 69 citations
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman et al.ICML 2024 · 45 citations
- Can Forward Gradient Match Backpropagation?Louis Fournier, Stéphane Rivaud, Eugene Belilovsky, Michael Eickenberg et al.ICML 2023 · 33 citations
Related papers
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li et al.ICCV 2025 · 3 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank StructuresYiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan et al.ICLR 2025
- FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale SpeedSizhe Dang, yangyangGuo, Yanjun Zhao, Xiaodong Zheng et al.ICLR 2026 · 16 citations
