Memory Savings at What Cost? A Study of Alternatives to Backpropagation
Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan
摘要
Forward-mode automatic differentiation (FMAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagationfree alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing. We present a unified theoretical and empirical comparison of BP, checkpointed BP, FMAD, and ZO for LLM and vision-language model training, showing that while FMAD and ZO reduce activation memory, they trade memory for higher computational cost and longer wallclock time to convergence, resulting in lower accuracy and slower training, especially under constrained perturbation budgets. Across models, BP with checkpointing outperforms FMAD and ZO variants, including variance-reduced methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence, and 3.8× fewer computations at comparable memory usage, while also revealing instability-related failure modes in FMAD and ZO. Overall, our results correct a one-sided benchmarking narrative by showing that memoryefficient methods entail fundamentally different trade-offs, and that ignoring these distinctions has led to misleading conclusions about LLM optimization in prior work. Our source code is available at https://github.com/Astuary/ Gradient_Estimation_Methods . This paper addresses the above-mentioned limitations through a comprehensive study of BP, FMAD, and ZO approaches in the context of LLM training and fine-tuning. We first outline the expected trade-offs among convergence behavior, memory consumption, and computational cost as functions of model dimensionality d and the number
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li 等ICML 2024 · 被引用 134 次
- Efficient Combination of Rematerialization and Offloading for Training DNNsOlivier Beaumont, Lionel Eyraud-Dubois, Alena ShilovaNeurIPS 2021 · 被引用 69 次
- Variance-reduced Zeroth-Order Methods for Fine-Tuning Language ModelsTanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman 等ICML 2024 · 被引用 45 次
- Can Forward Gradient Match Backpropagation?Louis Fournier, Stéphane Rivaud, Eugene Belilovsky, Michael Eickenberg 等ICML 2023 · 被引用 33 次
相关 Paper
- MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language ModelsYuezhang Peng, Yuxin Liu, Fei Wen, Xie ChenEMNLP 2025
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesZiming Yu, Pan Zhou, Sike Wang, Jia Li 等ICCV 2025 · 被引用 3 次
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank StructuresYiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan 等ICLR 2025
- FZOO: Fast Zeroth-Order Optimizer for Fine‑Tuning Large Language Models towards Adam‑Scale SpeedSizhe Dang, yangyangGuo, Yanjun Zhao, Xiaodong Zheng 等ICLR 2026 · 被引用 16 次
