VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity Control
Zi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei Wang
Abstract
As the model size of pre-trained language models (PLMs) grows rapidly, full fine-tuning becomes prohibitively expensive for model training and storage. In vision-and-language (VL), parameter-efficient tuning (PET) techniques are proposed to integrate modular modifications (e.g., Adapter and LoRA) into encoder-decoder PLMs. By tuning a small set of trainable parameters, these techniques perform on par with full fine-tuning. However, excessive modular modifications and neglecting the functionality gap between the encoders and decoders can lead to performance degradation, while existing PET techniques (e.g., VL-Adapter) overlook these critical issues. In this paper, we propose a Vision-and-Language Parameter-Efficient Tuning (VL-PET) framework to impose effective control over modular modifications via a novel granularity-controlled mechanism. Considering different granularity-controlled matrices generated by this mechanism, a variety of model-agnostic VL-PET modules can be instantiated from our framework for better efficiency and effective-ness trade-offs. We further propose lightweight PET module designs to enhance VL alignment and modeling for the encoders and maintain text generation for the decoders. Extensive experiments conducted on four image-text tasks and four video-text tasks demonstrate the efficiency, effectiveness and transferability of our VL-PET framework. In particular, our VL-PETlarge with lightweight PET module designs significantly outperforms VL-Adapter by 2.92% (3.41%) and LoRA by 3.37% (7.03%) with BART-base (T5-base) on image-text tasks. Furthermore, we validate the enhanced effect of employing our VL-PET designs on existing PET techniques, enabling them to achieve significant performance improvements. Our code is available at https://github.com/HenryHZY/VL-PET.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89e90497-f576-4f88-a6ff-da93fd2aa709Cited by top-tier papers9
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
- Memory-Space Visual Prompting for Efficient Vision-Language Fine-TuningShibo Jie, Yehui Tang, Ning Ding, Zhi-Hong Deng et al.ICML 2024 · 24 citations
- Personalized Federated Continual Learning via Multi-Granularity PromptHao Yu, Xin Yang, Xin Gao, Yan Kang et al.KDD 2024 · 12 citations
- LLaMA-Excitor: General Instruction Tuning via Indirect Feature InteractionBo Zou, Chao Yang, Yu Qiao, Chengbin Quan et al.CVPR 2024 · 5 citations
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental LearningHaojie Zhang, Yixiong Liang, Hulin Kuang, Lihui Cen et al.ACM MM 2025 · 2 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- VL-ADAPTER: Parameter-Efficient Transfer Learning for Vision-and-Language TasksYi-Lin Sung, Jaemin Cho, Mohit BansalCVPR 2022 · 22 citations
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi et al.ICLR 2024 · 11 citations
- VB-LoRA: Extreme Parameter Efficient Fine-Tuning with Vector BanksYang Li, Shaobo Han, Shihao JiNeurIPS 2024 · 61 citations
- Exploring Cross-Modal Flows for Few-Shot LearningZiqi Jiang, Yanghao Wang, Long ChenICLR 2026 · 6 citations
- HALoRA: Low-Rank Adaptation with Hierarchical Budget Allocation for Efficient Vision-Language AlignmentLetian Zhang, Guanghao Meng, Xudong Ren, Jinpeng WangAAAI 2026
