Parameter-efficient Tuning of Large-scale Multimodal Foundation Model
Haixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin, Jinan Sun, Shikun Zhang, Xiao Luo, Qi Tian
摘要
Driven by the progress of large-scale pre-training, parameter-efficient transfer learning has gained immense popularity across different subfields of Artificial Intelligence. The core is to adapt the model to downstream tasks with only a small set of parameters. Recently, researchers have leveraged such proven techniques in multimodal tasks and achieve promising results. However, two critical issues remain unresolved: how to further reduce the complexity with lightweight design and how to boost alignment between modalities under extremely low parameters. In this paper, we propose A graceful prompt framework for cross-modal transfer (Aurora) to overcome these challenges. Considering the redundancy in existing architectures, we first utilize the mode approximation to generate 0.1M trainable parameters to implement the multimodal parameter-efficient tuning, which explores the low intrinsic dimension with only 0.04% parameters of the pre-trained model. Then, for better modality alignment, we propose the Informative Context Enhancement and Gated Query Transformation module under extremely few parameters scenes. A thorough evaluation on six cross-modal benchmarks shows that it not only outperforms the state-of-the-art but even outperforms the full fine-tuning approach. Our code is available at: https://github.com/WillDreamer/Aurora .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- LION: Implicit Vision Prompt TuningHaixin Wang, Jianlong Chang, Yihang Zhai, Xiao Luo 等AAAI 2024 · 被引用 38 次
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 被引用 18 次
- Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language ModelsJiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li 等ACM MM 2024 · 被引用 5 次
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song 等NeurIPS 2025 · 被引用 5 次
- TPO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement LearningHaixin Wang, Hejie Cui, Chenwei Zhang, Jiahui Gao 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- CoPL: Parameter-Efficient Collaborative Prompt Learning for Audio-Visual TasksYihan Zhao, Wei Xi, Yuhang Cui, Gairui Bai 等ACM MM 2024 · 被引用 3 次
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen 等ICLR 2026 · 被引用 5 次
- Efficient Multimodal Fusion via Interactive PromptingYaowei Li, Ruijie Quan, Linchao Zhu, Yi YangCVPR 2023
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 等ICLR 2024 · 被引用 58 次
- Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningHao Zhao, Jie Fu, Zhaofeng HeEMNLP 2023 · 被引用 3 次
