Back Razor: Memory-Efficient Transfer Learning by Self-Sparsified Backpropagation
Ziyu Jiang, Xuxi Chen, Xueqin Huang, Xianzhi Du, Denny Zhou, Zhangyang Wang
摘要
Transfer learning from the model trained on large datasets to customized down-stream tasks has been widely used as the pre-trained model can greatly boost the generalizability. However, the increasing sizes of pre-trained models also lead to a prohibitively large memory footprints for downstream transferring, making them unaffordable for personal devices. Previous work recognizes the bottleneck of the footprint to be the activation, and hence proposes various solutions such as injecting specific lite modules. In this work, we present a novel memory-efficient transfer framework called Back Razor , that can be plug-and-play applied to any pre-trained network without changing its architecture. The key idea of Back Razor is asymmetric sparsifying : pruning the activation stored for back-propagation, while keeping the forward activation dense. It is based on the observation that the stored activation, that dominates the memory footprint, is only needed for back-propagation. Such asymmetric pruning avoids affecting the precision of forward computation, thus making more aggressive pruning possible. Furthermore, we conduct the theoretical analysis for the convergence rate of Back Razor, showing that under mild conditions, our method retains the similar convergence rate as vanilla SGD. Extensive transfer learning experiments on both Convolutional Neural Networks and Vision Transformers with classification, dense prediction, and language modeling tasks show that Back Razor could yield up to 97% sparsity , saving 9.2x memory usage, without losing accuracy. The code is available at: https://github.com/VITA-Group/BackRazor_Neurips22 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing BackpropagationYuchen Yang, Yingdong Shi, Cheems Wang, Xiantong Zhen 等ICML 2024 · 被引用 5 次
- Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised LearningSheng Li, Chao Wu, Ao Li, Yanzhi Wang 等ICLR 2024 · 被引用 4 次
- Sheared Backpropagation for Fine-Tuning Foundation ModelsZhiyuan Yu, Li Shen, Liang Ding, Xinmei Tian 等CVPR 2024 · 被引用 2 次
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun 等NeurIPS 2025 · 被引用 1 次
- Study of Training Dynamics for Memory-Constrained Fine-TuningAël Quélennec, Nour Hezbri, Pavlo Mozharovskyi, Van-Tam Nguyen 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
相关 Paper
- Resource- Efficient Transformer Pruning for Finetuning of Large ModelsFatih Ilhan, Gong Su, Selim Furkan Tekin, Tiansheng Huang 等CVPR 2024
- TinyTL: Reduce Memory, Not Parameters for Efficient On-Device LearningHan Cai, Chuang Gan, Ligeng Zhu, Song HanNeurIPS 2020 · 被引用 375 次
- RepNet: Efficient On-Device Learning via Feature ReprogrammingLi Yang, Adnan Siraj Rakin, Deliang FanCVPR 2022 · 被引用 18 次
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao 等NeurIPS 2023 · 被引用 293 次
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia 等NeurIPS 2023 · 被引用 18 次
