PELA: Learning Parameter-Efficient Models with Low-Rank Approximation
Yangyang Guo, Guangzhi Wang, Mohan S. Kankanhalli
摘要
Applying a pre-trained large model to downstream tasks is prohibitive under resource-constrained conditions. Recent dominant approaches for addressing efficiency issues involve adding a few learnable parameters to the fixed backbone model. This strategy, however, leads to more challenges in loading large models for downstream finetuning with limited resources. In this paper, we propose a novel method for increasing the parameter efficiency of pretrained models by introducing an intermediate pre-training stage. To this end, we first employ low-rank approximation to compress the original large model and then devise a feature distillation module and a weight perturbation regularization module. These modules are specifically designed to enhance the low-rank model. In particular, we update only the low-rank model while freezing the backbone parameters during pre-training. This allows for direct and efficient utilization of the low-rank model for downstream fine-tuning tasks. The proposed method achieves both efficiencies in terms of required parameters and computation time while maintaining comparable results with minimal modifications to the backbone architecture. Specifically, when applied to three vision-only and one vision-language Transformer models, our approach often demonstrates a merely ∼0.6 point decrease in performance while reducing the original parameter size by 1/3 to 2/3. We release our code at link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Lark: Low-Rank Updates After Knowledge Localization for Few-Shot Class-Incremental LearningJinxin Shi, Jiabao Zhao, Yifan Yang, Xingjiao Wu 等ICCV 2025 · 被引用 2 次
- S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum DomainBaoquan Zhang, Zhehao Yu, Lisai Zhang, Kenghong Lin 等CVPR 2026 · 被引用 1 次
- SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image GenerationLeigang Qu, Haochuan Li, Wenjie Wang, Xiang Liu 等CVPR 2025
- Adaptive Nonlinear Compression for Large Foundation ModelsLiang Xu, Shufan Shen, Qingming Huang, Yao Zhu 等ICLR 2026
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
相关 Paper
- HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language ModelsQiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang 等ICLR 2025
- WeGeFT: Weight‑Generative Fine-Tuning for Multi-Faceted Efficient Adaptation of Large ModelsChinmay Savadikar, Xi Song, Tianfu WuICML 2025
- Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design ApproachWei Dong, Xing Zhang, Bihui Chen, Dawei Yan 等CVPR 2024
- Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionXingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua 等EMNLP 2024 · 被引用 2 次
- Learning Global Controller in Latent Space for Parameter-Efficient Fine-TuningZeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong 等ACL 2024
