EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
Hsi-Che Lin, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank Wang
摘要
Open-source foundation models have seen rapid adoption and development, enabling powerful general-purpose capabilities across diverse domains. However, fine-tuning large foundation models for domain-specific or personalized tasks remains prohibitively expensive for most users due to the significant memory overhead beyond that of inference. We introduce EMLoC, an Emulator-based Memory-efficient fine-tuning framework with LoRA Correction, which enables model fine-tuning within the same memory budget required for inference. EMLoC constructs a task-specific light-weight emulator using activation-aware singular value decomposition (SVD) on a small downstream calibration set. Fine-tuning then is performed on this lightweight emulator via LoRA. To tackle the misalignment between the original model and the compressed emulator, we propose a novel compensation algorithm to correct the fine-tuned LoRA module, which thus can be merged into the original model for inference. EMLoC supports flexible compression ratios and standard training pipelines, making it adaptable to a wide range of applications. Extensive experiments demonstrate that EMLoC outperforms other baselines across multiple datasets and modalities. Moreover, without quantization, EMLoC enables fine-tuning of a 38B model, which originally required 95GB of memory, on a single 24GB consumer GPU-bringing efficient and practical model adaptation to individual users. Project Page: hsi-che-lin.github.io/EMLoC Introduction General-purpose foundation models have demonstrated impressive zero-shot capabilities across a wide range of benchmarks [2, 6, 8, 11, 38] . For real-world deployment, such as domain-specific tasks [22, 46] or personalized user behavior [28, 33] , further customization by fine-tuning is still required. However, fine-tuning typically incurs significantly more memory overhead than inference [30] . Consequently, if users have a fixed amount of available computing resources, they will be forced to choose between two unfavorable options. First, they can use a small model that fits within their memory budget for fine-tuning as shown in Fig. 1(a) , but this sacrifices the emergent capabilities [41] of larger models and underutilizes hardware during inference. Alternatively, they can opt for a large model that fully utilizes resources during inference but exceeds memory limits for fine-tuning as shown in Fig. 1 (b), making user-specific adaptation infeasible and potentially limiting performance in specialized applications. This paper addresses a central research question: Is it possible to design a fine-tuning strategy such that users can fine-tune a model under the same memory budget as inference? The memory cost of fine-tuning can be broadly attributed to three components: optimizer states, intermediate activations, and the model parameters themselves, as marked in Fig. 1 with different colors. Initial efforts to reduce the memory usage of fine-tuning concentrated on the first two components. The first component, optimizer states, stores auxiliary information such as momentum and variance in the Adam optimizer [16] for each trainable parameter. This overhead can be 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
相关 Paper
- Text-to-LoRA: Instant Transformer AdaptionRujikorn Charakorn, Edoardo Cetin, Yujin Tang, Robert Tjarko LangeICML 2025
- Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt AdaptationAbhinav Jain, Swarat Chaudhuri, Thomas W. Reps, Christopher M. JermaineNeurIPS 2024 · 被引用 10 次
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point ArithmeticKanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon 等NeurIPS 2025 · 被引用 1 次
- Parameter Efficient Fine-tuning via Explained Variance AdaptationFabian Paischer, Lukas Hauzenberger, Thomas Schmied, Benedikt Alkin 等NeurIPS 2025 · 被引用 25 次
- HiFT: A Hierarchical Full Parameter Fine-Tuning StrategyYongkang Liu, Yiqun Zhang, Qian Li, Tong Liu 等EMNLP 2024 · 被引用 7 次
