Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropagation
Ziyu Jiang, Xuxi Chen, Xueqin Huang, Xianzhi Du, Denny Zhou, Zhangyang Wang
Abstract
Transfer learning from the model trained on large datasets to customized down-stream tasks has been widely used as the pre-trained model can greatly boost the generalizability. However, the increasing sizes of pre-trained models also lead to a prohibitively large memory footprints for downstream transferring, making them unaffordable for personal devices. Previous work recognizes the bottleneck of the footprint to be the activation, and hence proposes various solutions such as injecting specific lite modules. In this work, we present a novel memory-efficient transfer framework called Back Razor , that can be plug-and-play applied to any pre-trained network without changing its architecture. The key idea of Back Razor is asymmetric sparsifying : pruning the activation stored for back-propagation, while keeping the forward activation dense. It is based on the observation that the stored activation, that dominates the memory footprint, is only needed for back-propagation. Such asymmetric pruning avoids affecting the precision of forward computation, thus making more aggressive pruning possible. Furthermore, we conduct the theoretical analysis for the convergence rate of Back Razor, showing that under mild conditions, our method retains the similar convergence rate as vanilla SGD. Extensive transfer learning experiments on both Convolutional Neural Networks and Vision Transformers with classification, dense prediction, and language modeling tasks show that Back Razor could yield up to 97% sparsity , saving 9.2x memory usage, without losing accuracy. The code is available at: https://github.com/VITA-Group/BackRazor_Neurips22 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing BackpropagationYuchen Yang, Yingdong Shi, Cheems Wang, Xiantong Zhen et al.ICML 2024 · 5 citations
- Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningSheng Li, Chao Wu, Ao Li, Yanzhi Wang et al.ICLR 2024 · 4 citations
- Sheared Backpropagation for Fine-Tuning Foundation ModelsZhiyuan Yu, Li Shen, Liang Ding, Xinmei Tian et al.CVPR 2024 · 2 citations
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun et al.NeurIPS 2025 · 1 citation
- Study of Training Dynamics for Memory-Constrained Fine-TuningAël Quélennec, Nour Hezbri, Pavlo Mozharovskyi, Van-Tam Nguyen et al.ICLR 2026 · 1 citation
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
Related papers
- Resource- Efficient Transformer Pruning for Finetuning of Large ModelsFatih Ilhan, Gong Su, Selim Furkan Tekin, Tiansheng Huang et al.CVPR 2024
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 375 citations
- RepNet: Efficient On-Device Learning via Feature ReprogrammingLi Yang, Adnan Siraj Rakin, Deliang FanCVPR 2022 · 18 citations
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao et al.NeurIPS 2023 · 293 citations
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia et al.NeurIPS 2023 · 18 citations
