From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel Control
Chi Zhang, Lianhai Ren, Jingpu Cheng, Qianxiao Li
Abstract
The LoRA method has achieved notable success in reducing GPU memory usage by applying lowrank updates to weight matrices. Yet, one simple question remains: can we push this reduction even further? Furthermore, is it possible to achieve this while reducing computation time and preserving performance? Answering these questions requires moving beyond the conventional weight-centric approach. In this paper, we present a state-based fine-tuning framework that shifts the focus from weight adaptation to optimizing forward states, with LoRA acting as a special example. Specifically, state-based tuning introduces parameterized perturbations to the states within the computational graph, allowing us to control states across an entire residual block. A key advantage of this approach is the potential to avoid storing large intermediate states in models like transformers. Empirical results across multiple architectures-including ViT, RoBERTa, LLaMA2-7B, and LLaMA3-8B-show that our method further reduces memory consumption and computation time while preserving performance. As a result of memory reduction, we explore the feasibility to train 7B/8B models on consumer-level GPUs like Nvidia 3090, without model quantization. The code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Machine Unlearning under Retain–Forget EntanglementJingpu Cheng, Ping Liu, Qianxiao Li, CHI ZHANGICLR 2026 · 11 citations
- Closed-Form Concept Erasure via Double ProjectionsChi Zhang, Jingpu Cheng, Zhixian Wang, Ping LiuCVPR 2026 · 5 citations
- A unified framework for establishing the universal approximation of transformer-type architecturesJingpu Cheng, Ting Lin, Zuowei Shen, Qianxiao LiNeurIPS 2025 · 2 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 308 citations
Related papers
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- NOLA: Compressing LoRA using Linear Combination of Random BasisSoroush Abbasi Koohpayegani, Navaneet K. L., Parsa Nooralinejad, Soheil Kolouri et al.ICLR 2024 · 33 citations
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel et al.ICLR 2025
- Gradient Weight-normalized Low-rank Projection for Efficient LLM TrainingJia-Hong Huang, Yixian Shen, Hongyi Zhu, Stevan Rudinac et al.AAAI 2025 · 1 citation
