Memory-Efficient Fine-Tuning of Transformers via Token Selection
Antoine Simoulin, Namyong Park, Xiaoyi Liu, Grey Yang
Abstract
Fine-tuning provides an effective means to specialize pre-trained models for various downstream tasks. However, fine-tuning often incurs high memory overhead, especially for large transformer-based models, such as LLMs. While existing methods may reduce certain parts of the memory required for fine-tuning, they still require caching all intermediate activations computed in the forward pass to update weights during the backward pass. In this work, we develop TOKENTUNE, a method to reduce memory usage, specifically the memory to store intermediate activations, in the finetuning of transformer-based models. During the backward pass, TOKENTUNE approximates the gradient computation by backpropagating through just a subset of input tokens. Thus, with TOKENTUNE, only a subset of intermediate activations are cached during the forward pass. Also, TOKENTUNE can be easily combined with existing methods like LoRA, further reducing the memory cost. We evaluate our approach on pre-trained transformer models with up to billions of parameters, considering the performance on multiple downstream tasks such as text classification and question answering in a few-shot learning setup. Overall, TOKENTUNE achieves performance on par with full fine-tuning or representative memoryefficient fine-tuning methods, while greatly reducing the memory footprint, especially when combined with other methods with complementary memory reduction mechanisms. We hope that our approach will facilitate the finetuning of large transformers, in specializing them for specific domains or co-training them with other neural components from a larger system. Our code is available at https://github. com/facebookresearch/tokentune .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 712ca40f-62b3-41b8-a834-0a45d97a970aCited by top-tier papers4
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuningXiaohan Qin, Victor Wang, Ning Liao, Cancheng Zhang et al.ICLR 2026 · 3 citations
- TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token DitchingRunjia Zeng, Qifan Wang, Qiang Guan, Ruixiang Tang et al.ICLR 2026 · 1 citation
- TokenDrop: Token-Level Importance-Aware Backward Propagation Skipping for Efficient LLM Fine-TuningBeomseok Kim, Sol Namkung, Dongsuk JeonICML 2026
- Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language ModelsYeachan Kim, SangKeun LeeACL 2025
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 22 citations
- Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language ModelsZhengxin Zhang, Dan Zhao, Xupeng Miao, Gabriele Oliaro et al.ACL 2024
- From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel ControlChi Zhang, Lianhai Ren, Jingpu Cheng, Qianxiao LiICML 2025
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMsKairun Zhang, Haoyu Li, Yanjun Zhao, Yifan Sun et al.ICML 2026 · 1 citation
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 347 citations
