Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Yang You, Guiming Xie, Xuejian Gong, Kunlong Zhou
摘要
Large Language Models (LLMs) have significantly advanced natural language processing with exceptional task generalization capabilities. Low-Rank Adaption (LoRA) offers a cost-effective fine-tuning solution, freezing the original model parameters and training only lightweight, low-rank adapter matrices. However, the memory footprint of LoRA is largely dominated by the original model parameters. To mitigate this, we propose LORAM, a memory-efficient LoRA training scheme founded on the intuition that many neurons in over-parameterized LLMs have low training utility but are essential for inference. LORAM presents a unique twist: it trains on a pruned (small) model to obtain pruned low-rank matrices, which are then recovered and utilized with the original (large) model for inference. Additionally, minimal-cost continual pre-training, performed by the model publishers in advance, aligns the knowledge discrepancy between pruned and original models. Our extensive experiments demonstrate the efficacy of LORAM across various pruning strategies and downstream tasks. For a model with 70 billion parameters, LORAM enables training on a GPU with only 20G HBM, replacing an A100-80G GPU for LoRA training and 15 GPUs for full fine-tuning. Specifically, QLORAM implemented by structured pruning combined with 4-bit quantization, for LLaMA-3.1-70B (LLaMA-2-70B), reduces the parameter storage cost that dominates the memory usage in low-rank matrix training by 15.81× (16.95×), while achieving dominant performance gains over both the original LLaMA-3.1-70B (LLaMA-2-70B) and LoRA-trained LLaMA-3.1-8B (LLaMA-2-13B). Code is available at https://github.com/junzhang-zj/LoRAM . * Work done during an internship at OPPO AI Center. † Huan Li and Lidan Shou are the corresponding authors. 1 Training is performed on one sample with a length of 4K using BF16 mixed precision with the Adam optimizer, incorporating gradient checkpointing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMsZhanda Zhu, Qidong Su, Yaoyao Ding, Kevin Song 等EuroSys 2026 · 被引用 2 次
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual SparsitySamir Khaki, Xiuyu Li, Junxian Guo, Ligeng Zhu 等ICML 2025
- EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA CorrectionHsi-Che Lin, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank WangNeurIPS 2025
- Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base ModelsNaibin Gu, Peng Fu, Xiyu Liu, Ke Ma 等ACL 2025
它引用的顶会 Paper36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- LQ-LoRA: Low-rank plus Quantized Matrix Decomposition for Efficient Language Model FinetuningHan Guo, Philip Greengard, Eric P. Xing, Yoon KimICLR 2024 · 被引用 94 次
- Taming Momentum: Rethinking Optimizer States Through Low-Rank ApproximationZhengbo Wang, Jian Liang, Ran He, Zilei Wang 等ICLR 2026 · 被引用 3 次
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 等ICLR 2025
- Parameter and Memory Efficient Pretraining via Low-rank Riemannian OptimizationZhanfeng Mo, Long-Kai Huang, Sinno Jialin PanICLR 2025
