Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-Tuning
Baohao Liao, Shaomu Tan, Christof Monz
Abstract
Parameter-efficient fine-tuning (PEFT) of pre-trained language models (PLMs) has emerged as a highly successful approach, with training only a small number of parameters without sacrificing performance and becoming the de-facto learning paradigm with the increasing size of PLMs. However, existing PEFT methods are not memory-efficient, because they still require caching most of the intermediate activations for the gradient calculation, akin to fine-tuning. One effective way to reduce the activation memory is to apply a reversible model, so the intermediate activations are not necessary to be cached and can be recomputed. Nevertheless, modifying a PLM to its reversible variant is not straightforward, since the reversible model has a distinct architecture from the currently released PLMs. In this paper, we first investigate what is a key factor for the success of existing PEFT methods, and realize that it's essential to preserve the PLM's starting point when initializing a PEFT method. With this finding, we propose memory-efficient fine-tuning (MEFT) that inserts adapters into a PLM, preserving the PLM's starting point and making it reversible without additional pre-training. We evaluate MEFT on the GLUE benchmark and five question-answering tasks with various backbones, BERT, RoBERTa, BART and OPT. MEFT significantly reduces the activation memory up to 84% of full fine-tuning with a negligible amount of trainable parameters. Moreover, MEFT achieves the same score on GLUE and a comparable score on the question-answering tasks as full fine-tuning. A similar finding is also observed for the image classification task. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaaa64aa-0286-43d4-a9dd-8e186cd7db4dCited by top-tier papers12
- 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and ComposabilityBaohao Liao, Christof MonzNeurIPS 2024 · 14 citations
- MEFT: Memory-Efficient Fine-Tuning through Sparse AdapterJitai Hao, Weiwei Sun, Xin Xin, Qi Meng et al.ACL 2024 · 4 citations
- DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language ModelsYanming Liu, Xinyue Peng, Yuwei Zhang, Xiaolan Ke et al.AAAI 2025 · 4 citations
- QStore: Quantization-Aware Compressed Model StorageRaunak Shah, Zhaoheng Li, Yongjoo ParkVLDB 2026 · 3 citations
- Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and InferenceSiyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li et al.ACL 2025 · 3 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Parameter-Efficient Fine-Tuning without Introducing New LatencyBaohao Liao, Yan Meng, Christof MonzACL 2023 · 26 citations
- PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from AttentionHaonan Wang, Brian K Chen, Siquan Li, Liang Xinhe et al.ICLR 2026 · 5 citations
- From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in TransformersBharat Runwal, Tejaswini Pedapati, Pin-Yu ChenAAAI 2025 · 10 citations
- Parameter-efficient Tuning for Large Language Model without Calculating Its GradientsFeihu Jin, Jiajun Zhang, Chengqing ZongEMNLP 2023 · 2 citations
- AdaMix: Mixture-of-Adaptations for Parameter-efficient Model TuningYaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu et al.EMNLP 2022 · 65 citations
