VioLET: Vision-Language Efficient Tuning with Collaborative Multi-modal Gradients
Yaoming Wang, Yuchen Liu, Xiaopeng Zhang, Jin Li, Bowen Shi, Chenglin Li, Wenrui Dai, Hongkai Xiong, Qi Tian
Abstract
Parameter-Efficient Tuning (PET) has emerged as a leading advancement in both Natural Language Processing and Computer Vision, enabling efficient accommodation of downstream tasks without costly fine-tuning. However, most existing PET approaches are limited to uni-modal tuning, even for vision-language models like CLIP. We investigate this limitation and demonstrate that simultaneous tuning of the two modalities in such models leads to multi-modal forgetting and catastrophic performance degradation, particularly when generalizing to new classes. To address this issue, we propose a novel PET approach called VioLET (Vision Language Efficient Tuning) that utilizes collaborative multi-modal gradients to unlock the full potential of both modalities. Specifically, we incorporate an additional visual encoder without learnable parameters and use these two visual encoders to compute the gradients of the context parameters separately. When conflicts arise, we replace the original gradient with an orthogonal gradient. Extensive experiments are conducted on few-shot recognition and unseen class generalization tasks using ResNet-50 or ViT/B-16 as the backbone. VioLET consistently outperforms several state-of-the-art methods on 11 datasets, showcasing its superiority over existing PET approaches. The code is available at https://github.com/Wang-Yaoming/VioLET.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c6f8f918-79ef-4bc6-8b2a-aaedd0eca5c6Cited by top-tier papers1
Ask how each one uses itRelated papers
- Adaptive Parameter Selection for Tuning Vision-Language ModelsYi Zhang, Yi-Xuan Deng, Meng-Hao Guo, Shi-Min HuCVPR 2025
- FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsZhengqin Xu, Zelin Peng, Xiaokang Yang, Wei ShenAAAI 2025 · 3 citations
- Fed-Duet: Dual Expert-Orchestrated Framework for Continual Federated Vision-Language LearningTao Guo, Junwei Chen, Laizhong CuiICLR 2026
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan et al.AAAI 2024 · 102 citations
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsSanjoy Chowdhury, Sayan Nag, Dinesh ManochaEMNLP 2023 · 17 citations
