Sheared Backpropagation for Fine-Tuning Foundation Models
Zhiyuan Yu, Li Shen, Liang Ding, Xinmei Tian, Yixin Chen, Dacheng Tao
Abstract
Fine-tuning is the process of extending the training of pre-trained models on specific target tasks, thereby significantly enhancing their performance across various applications. However, fine-tuning often demands large memory consumption, posing a challenge for low-memory devices that some previous memory-efficient fine-tuning methods attempted to mitigate by pruning activations for gradient computation, albeit at the cost of significant computational overhead from the pruning processes during training. To address these challenges, we introduce PreBackRazor; a novel activation pruning scheme offering both computational and memory efficiency through a sparsified back-propagation strategy, which judiciously avoids unnecessary activation pruning and storage and gradient computation. Before activation pruning, our approach samples a probability of selecting a portion of parameters to freeze, utilizing a bandit method for updates to prioritize impactful gradients on convergence. During the feed-forward pass, each model layer adjusts adaptively based on parameter activation status, obviating the need for sparsification and storage of redundant activations for subsequent backpropagation. Benchmarking on fine-tuning foundation models, our approach maintains baseline accuracy across diverse tasks, yielding over 20% speedup and around 10% memory reduction. Moreover, integrating with an advanced CUDA kernel achieves up to 60% speedup without extra memory costs or accuracy loss, significantly enhancing the efficiency of fine-tuning foundation models on memory-constrained devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e47cd0b4-1ea9-4b5d-91e4-5fb490476745Cited by top-tier papers2
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun et al.NeurIPS 2025 · 1 citation
- Subspace Optimization for Large Language Models with Convergence GuaranteesYutong He, Pengrui Li, Yipeng Hu, Chuyan Chen et al.ICML 2025
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
Related papers
- Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified BackpropagationZiyu Jiang, Xuxi Chen, Xueqin Huang, Xianzhi Du et al.NeurIPS 2022 · 25 citations
- Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language ModelsYeachan Kim, SangKeun LeeACL 2025
- PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient TrainingYanyi Li, Yimu Zhang, Cong FangICML 2026
- DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank DistributionYulong Mao, Kaiyu Huang, Changhao Guan, Ganglin Bao et al.ACL 2024 · 15 citations
- DRAFT: Decoupling Backpropagation from Pre-trained Backbone for Efficient Transformer Fine-Tuning on EdgeZhirui Huang, Shiwei Liu, Haozhe Zhu, Qi Liu et al.DAC 2025
