CoPL: Parameter-Efficient Collaborative Prompt Learning for Audio-Visual Tasks
Yihan Zhao, Wei Xi, Yuhang Cui, Gairui Bai, Xinhui Liu, Jizhong Zhao
Abstract
Parameter-Efficient Fine Tuning (PEFT) has been demonstrated to be effective and efficient for transferring foundation models to downstream tasks. Transferring pretrained uni-modal models to multi-modal downstream tasks helps alleviate substantial computational costs for retraining multi-modal models. However, existing approaches primarily focus on multi-modal fusion, while neglecting the modal-specific fine-tuning, which is also crucial for multi-modal tasks. To this end, we propose parameter-efficient Collaborative Prompt Learning (CoPL) to fine-tune both uni-modal and multi-modal features. Specifically, the collaborative prompts consist of modal-specific prompts and modal-interaction prompts. The modal-specific prompts are tailored for fine-tuning each modality, while the modal-interaction prompts are customized to explore inter-modality association. Furthermore, prompt bank-based mutual coupling is introduced to extract instance-level features, further enhancing the model's generalization ability. Extensive experimental results demonstrate that our approach achieves comparable or higher performance on various audio-visual downstream tasks while utilizing approximately 1% extra trainable parameters.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Efficient Multimodal Fusion via Interactive PromptingYaowei Li, Ruijie Quan, Linchao Zhu, Yi YangCVPR 2023
- FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsZhengqin Xu, Zelin Peng, Xiaokang Yang, Wei ShenAAAI 2025 · 3 citations
- DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningZhengxiang Shi, Aldo LipaniICLR 2024 · 46 citations
- CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream TasksWish Suharitdamrong, Tony Alex, Muhammad Awais, Sara AtitoICML 2026
- M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningTaowen Wang, Yiyang Liu, James Liang, Junhan Zhao et al.EMNLP 2024 · 31 citations
