Memory-Efficient Fine-Tuning of Transformers via Token Selection
Antoine Simoulin, Namyong Park, Xiaoyi Liu, Grey Yang
摘要
Fine-tuning provides an effective means to specialize pre-trained models for various downstream tasks. However, fine-tuning often incurs high memory overhead, especially for large transformer-based models, such as LLMs. While existing methods may reduce certain parts of the memory required for fine-tuning, they still require caching all intermediate activations computed in the forward pass to update weights during the backward pass. In this work, we develop TOKENTUNE, a method to reduce memory usage, specifically the memory to store intermediate activations, in the finetuning of transformer-based models. During the backward pass, TOKENTUNE approximates the gradient computation by backpropagating through just a subset of input tokens. Thus, with TOKENTUNE, only a subset of intermediate activations are cached during the forward pass. Also, TOKENTUNE can be easily combined with existing methods like LoRA, further reducing the memory cost. We evaluate our approach on pre-trained transformer models with up to billions of parameters, considering the performance on multiple downstream tasks such as text classification and question answering in a few-shot learning setup. Overall, TOKENTUNE achieves performance on par with full fine-tuning or representative memoryefficient fine-tuning methods, while greatly reducing the memory footprint, especially when combined with other methods with complementary memory reduction mechanisms. We hope that our approach will facilitate the finetuning of large transformers, in specializing them for specific domains or co-training them with other neural components from a larger system. Our code is available at https://github. com/facebookresearch/tokentune .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuningXiaohan Qin, Victor Wang, Ning Liao, Cancheng Zhang 等ICLR 2026 · 被引用 3 次
- TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token DitchingRunjia Zeng, Qifan Wang, Qiang Guan, Ruixiang Tang 等ICLR 2026 · 被引用 1 次
- TokenDrop: Token-Level Importance-Aware Backward Propagation Skipping for Efficient LLM Fine-TuningBeomseok Kim, Sol Namkung, Dongsuk JeonICML 2026
- Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language ModelsYeachan Kim, SangKeun LeeACL 2025
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 被引用 22 次
- Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language ModelsZhengxin Zhang, Dan Zhao, Xupeng Miao, Gabriele Oliaro 等ACL 2024
- From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel ControlChi Zhang, Lianhai Ren, Jingpu Cheng, Qianxiao LiICML 2025
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMsKairun Zhang, Haoyu Li, Yanjun Zhao, Yifan Sun 等ICML 2026 · 被引用 1 次
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 被引用 347 次
