PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
Sunghyeon Woo, Sol Namkung, Sunwoo Lee, Inho Jeong, Beomseok Kim, Dongsuk Jeon
Abstract
Prior parameter-efficient fine-tuning (PEFT) algorithms reduce memory usage and computational costs of fine-tuning large neural network models by training only a few additional adapter parameters, rather than the entire model. However, the reduction in computational costs due to PEFT does not necessarily translate to a reduction in training time; although the computational costs of the adapter layers are much smaller than the pretrained layers, it is well known that those two types of layers are processed sequentially on GPUs, resulting in significant latency overhead. LoRA and its variants avoid this latency overhead by merging the low-rank adapter matrices with the pretrained weights during inference. However, those layers cannot be merged during training since the pretrained weights must remain frozen while the low-rank adapter matrices are updated continuously over the course of training. Furthermore, LoRA and its variants do not reduce activation memory, as the first low-rank adapter matrix still requires the input activations to the pretrained weights to compute weight gradients. To mitigate this issue, we propose Partial Connection Adaptation (PaCA), which fine-tunes randomly selected partial connections within the pretrained weights instead of introducing adapter layers in the model. PaCA not only enhances training speed by eliminating the time overhead due to the sequential processing of the adapter and pretrained layers but also reduces activation memory since only partial activations, rather than full activations, need to be stored for gradient computation. Compared to LoRA, PaCA reduces training time by 22% and total memory usage by 16%, while maintaining comparable accuracy across various fine-tuning scenarios, such as fine-tuning on the MMLU dataset and instruction tuning on the Oasst1 dataset. PaCA can also be combined with quantization, enabling the finetuning of large models such as LLaMA3.1-70B. In addition, PaCA enables training with 23% longer sequence and improves throughput by 16% on both NVIDIA A100 GPU and INTEL Gaudi2 HPU compared to LoRA. The code is available at https://github.com/WooSunghyeon/paca .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ICaRus: Identical Cache Reuse for Efficient Multi-Model InferenceSunghyeon Woo, Jaeeun Kil, Hoseung Kim, Minsub Kim et al.ICLR 2026 · 7 citations
- Study of Training Dynamics for Memory-Constrained Fine-TuningAël Quélennec, Nour Hezbri, Pavlo Mozharovskyi, Van-Tam Nguyen et al.ICLR 2026 · 1 citation
- S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum DomainBaoquan Zhang, Zhehao Yu, Lisai Zhang, Kenghong Lin et al.CVPR 2026 · 1 citation
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-TuningZhi Zhang, Yixian Shen, Congfeng Cao, Ekaterina ShutovaEMNLP 2025
- From Bottom to Top: Extending the Potential of Parameter Efficient Fine-TuningJihao Gu, Zelin Wang, Yibo Zhang, Ziji Zhang et al.EMNLP 2024 · 3 citations
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and OptimizationXujia Wang, Yunjia Qi, Bin XuEMNLP 2025
- VB-LoRA: Extreme Parameter Efficient Fine-Tuning with Vector BanksYang Li, Shaobo Han, Shihao JiNeurIPS 2024 · 61 citations
- AdaMix: Mixture-of-Adaptations for Parameter-efficient Model TuningYaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu et al.EMNLP 2022 · 65 citations
