PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches
Rana Muhammad Shahroz, Pingzhi Li, Sukwon Yun, Zhenyu Wang, Shahriar Nirjon, Chau-Wai Wong, Tianlong Chen
Abstract
As large language models (LLMs) increasingly shape the AI landscape, finetuning pretrained models has become more popular than it was in the pre-LLM era for achieving optimal performance in domain-specific tasks. However, pretrained LLMs such as ChatGPT are periodically evolved (i.e., model parameters are frequently updated), making it challenging for downstream users with limited resources to keep up with fine-tuning the newest LLMs for their domain application. Even though fine-tuning costs have nowadays been reduced thanks to the innovations in parameter-efficient fine-tuning such as low-rank adaptation (LoRA), not all downstream users have adequate computing for frequent personalization. Moreover, access to fine-tuning datasets, particularly in sensitive domains such as healthcare, can be time-restrictive, making it crucial to retain the knowledge encoded in earlier fine-tuned rounds for future adaptation. In this paper, we present PORTLLM, a training-free framework that (i) creates an initial lightweight model update patch to capture domain-specific knowledge, and (ii) allows a subsequent seamless plugging for the continual personalization of the evolved LLM at minimal cost. Our extensive experiments cover seven representative datasets, from easier question-answering tasks BoolQ, SST2 to harder reasoning tasks WinoGrande, GSM8K, and models including Mistral-7B, Llama2, Llama3.1, and Gemma2, validating the portability of our designed model patches and showcasing the effectiveness of our proposed framework. For instance, PORTLLM achieves comparable performance to LoRA fine-tuning with reductions of up to 12.2× in GPU memory usage. Finally, we provide theoretical justifications to understand the portability of our model update patches, which offers new insights into the theoretical dimension of LLMs' personalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f23e5c34-d743-4193-911f-eb985a4cd963Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- CAR-LoRA: Training Compression-Aware and Robust LoRA Adapters for Evolving LLMsRana Muhammad Shahroz Khan, Zhen Tan, Ruichen Zhang, Hua Wei et al.ICLR 2026
- Trans-LoRA: towards data-free Transferable Parameter Efficient FinetuningRunqian Wang, Soumya Ghosh, David D. Cox, Diego Antognini et al.NeurIPS 2024 · 15 citations
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-TuningXiao Han, Zimo Zhao, Wanyu Wang, Maolin Wang et al.NeurIPS 2025 · 4 citations
- Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language ModelsJun Zhang, Jue Wang, Huan Li, Lidan Shou et al.ICLR 2025
- LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA OptimizationJui-Nan Yen, Si Si, Zhao Meng, Felix X. Yu et al.ICLR 2025
