MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
Xiaojie Jin, Bowen Zhang, Weibo Gong, Kai Xu, Xueqing Deng, Peng Wang, Zhao Zhang, Xiaohui Shen, Jiashi Feng
Abstract
State-of-the-art video-text retrieval (VTR) methods typically involve fully fine-tuning a pre-trained model (e.g. CLIP) on specific datasets. However, this can result in significant storage costs in practical applications as a separate model per task must be stored. To address this issue, we present our pioneering work that enables parameterefficient VTR using a pre-trained model, with only a small number of tunable parameters during training. Towards this goal, we propose a new method dubbed Multimodal Video Adapter (MV-Adapter) for efficiently transferring the knowledge in the pre-trained CLIP from image-text to video-text. Specifically, MV-Adapter utilizes bottleneck structures in both video and text branches, along with two novel components. The first is a Temporal Adaptation Module that is incorporated in the video branch to introduce global and local temporal contexts. We also train weights calibrations to adjust to dynamic variations across frames. The second is Cross Modality Tying that generates weights for video/text branches through sharing cross modality factors, for better aligning between modalities. Thanks to above innovations, MV-Adapter can achieve comparable or better performance than standard full fine-tuning with negligible parameters overhead. Notably, MV-Adapter consistently outperforms various competing methods in V2T/T2V tasks with large margins on five widely used VTR benchmarks (MSR-VTT, MSVD, LSMDC, DiDemo, and Activi-tyNet). Codes will be released. * Equal contribution. Bowen Zhang did the work during an internship.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4c8a8b8-0936-49cb-86b6-59a9985a6c32Cited by top-tier papers9
- Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video RetrievalBangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun et al.ICCV 2025 · 5 citations
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow ModelsMinghao Yin, Yukang Cao, Kai HanNeurIPS 2025 · 5 citations
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video RetrievalDohwan Ko, Ji Soo Lee, Minhyuk Choi, Zihang Meng et al.ICCV 2025 · 4 citations
- Spatial Alignment and Temporal Matching Adapter for Video-Radar Remote Physiological MeasurementQian Liang, Ruixu Geng, Jinbo Chen, Haoyu Wang et al.ICCV 2025 · 3 citations
- SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video SegmentationClaudia Cuttano, Gabriele Trivigno, Gabriele Rosi, Carlo Masone et al.CVPR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
- MPT: Multi-grained Prompt Tuning for Text-Video RetrievalHaonan Zhang, Pengpeng Zeng, Lianli Gao, Jingkuan Song et al.ACM MM 2024 · 16 citations
- VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal RetrievalSiteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang et al.CVPR 2023
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 53 citations
- A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionMengmeng Wang, Jiazheng Xing, Boyuan Jiang, Jun Chen et al.AAAI 2024 · 13 citations
