Efficient Multimodal Fusion via Interactive Prompting
Yaowei Li, Ruijie Quan, Linchao Zhu, Yi Yang
Abstract
Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multimodal learning models constantly increases, leading to an urgent need to reduce the massive computational cost of finetuning these models for downstream tasks. In this paper, we propose an efficient and flexible multimodal fusion method, namely PMF, tailored for fusing unimodally pretrained transformers. Specifically, we first present a modular multimodal fusion framework that exhibits high flexibility and facilitates mutual interactions among different modalities. In addition, we disentangle vanilla prompts into three types in order to learn different optimizing objectives for multimodal learning. It is also worth noting that we propose to add prompt vectors only on the deep layers of the unimodal transformers, thus significantly reducing the training memory usage. Experiment results show that our proposed method achieves comparable performance to several other multimodal finetuning methods with less than 3% trainable parameters and up to 66% saving of training memory usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Federated Text-driven Prompt Generation for Vision-Language ModelsChen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh et al.ICLR 2024 · 33 citations
- Learning Visual Prompt for Gait RecognitionKang Ma, Ying Fu, Chunshui Cao, Saihui Hou et al.CVPR 2024 · 24 citations
- Dango: A Mixed-Initiative Data Wrangling System using Large Language ModelWei-Hao Chen, Weixi Tong, Amanda Case, Tianyi ZhangCHI 2025 · 19 citations
- On Disentanglement of Asymmetrical Knowledge Transfer for Modality-Task Agnostic Federated LearningJiayi Chen, Aidong ZhangAAAI 2024 · 18 citations
- Psychometry: An Omnifit Model for Image Reconstruction from Human Brain ActivityRuijie Quan, Wenguan Wang, Zhibo Tian, Fan Ma et al.CVPR 2024 · 18 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- CoPL: Parameter-Efficient Collaborative Prompt Learning for Audio-Visual TasksYihan Zhao, Wei Xi, Yuhang Cui, Gairui Bai et al.ACM MM 2024 · 3 citations
- EPT: Efficient Prompt Tuning by Multi-Space Projection and Prompt FusionPengxiang Lan, Enneng Yang, Yuting Liu, Guibing Guo et al.AAAI 2025 · 4 citations
- Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelHaixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin et al.NeurIPS 2023 · 45 citations
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
- Multimodal Prompting with Missing Modalities for Visual RecognitionYi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, Chen-Yu LeeCVPR 2023
