UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
Haoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu, Wei Zhan, Masayoshi Tomizuka, Mingyu Ding
摘要
Large-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the conventional full fine-tuning paradigm becomes impractical due to heavy computational and storage costs. This paper proposes UniAdapter, which unifies unimodal and multimodal adapters for parameter-efficient cross-modal adaptation on pre-trained vision-language models. Specifically, adapters are distributed to different modalities and their interactions, with the total number of tunable parameters reduced by partial weight sharing. The unified and knowledge-sharing design enables efficient adaptation to various downstream tasks with powerful cross-modal representations, requiring only 1.0%-2.0% tunable parameters of the pre-trained model. Extensive experiments on 6 crossmodal downstream benchmarks (including video-text retrieval, image-text retrieval, VideoQA, and VQA) show that in most cases, UniAdapter not only outperforms the state-of-the-arts, but even surpasses the full fine-tuning strategy. Notably, on the MSRVTT retrieval task, UniAdapter achieves 49.7% recall@1 with only 2.2% tunable model parameters, outperforming the latest competitors by 2.0%. The code and models are available at https://github.com/UniAdapter/UniAdapter .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Unified Coarse-to-Fine Alignment for Video-Text RetrievalZiyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius 等ICCV 2023 · 被引用 90 次
- Parameter-efficient Tuning of Large-scale Multimodal Foundation ModelHaixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin 等NeurIPS 2023 · 被引用 45 次
- VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity ControlZi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei WangICCV 2023 · 被引用 25 次
- Global Minimizers of Sigmoid Contrastive LossKiril Bangachev, Guy Bresler, Iliyas Noman, Yury PolyanskiyNeurIPS 2025 · 被引用 5 次
- DitHub: A Modular Framework for Incremental Open-Vocabulary Object DetectionChiara Cappellino, Gianluca Mancusi, Matteo Mosconi, Angelo Porrello 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- VL-ADAPTER: Parameter-Efficient Transfer Learning for Vision-and-Language TasksYi-Lin Sung, Jaemin Cho, Mohit BansalCVPR 2022 · 被引用 22 次
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
- MV-Adapter: Multimodal Video Transfer Learning for Video Text RetrievalXiaojie Jin, Bowen Zhang, Weibo Gong, Kai Xu 等CVPR 2024
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen 等ICLR 2026 · 被引用 5 次
- Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video PretrainingZihao Liu, Xiaoyu Wu, Shengjin Wang, Jiayao QianACM MM 2024 · 被引用 1 次
