MPT: Multi-grained Prompt Tuning for Text-Video Retrieval
Haonan Zhang, Pengpeng Zeng, Lianli Gao, Jingkuan Song, Heng Tao Shen
摘要
Recently, significant advancements have been made in supporting text-video retrieval by transferring large-scale image-text pre-training models through model adaptation, i.e., full fine-tuning, or prompt tuning, a parameter-efficient fine-tuning strategy. While full fine-tuning involves high computational costs, particularly with increasing model size, prompt tuning offers greater flexibility and efficiency by adjusting only a few learnable parameters. However, current prompt tuning methods rely on coarse visual and textual cues for text-video retrieval task, neglecting the domain-specific features when performing the adaptation. This approach may lead to sub-optimal performance due to the incorporation of irrelevant and indiscriminate knowledge. To address such an issue, we present a Multi-grained Prompt Tuning (MPT) for text-video retrieval, that designs a variety of specific prompts to effectively explore semantic interaction across different modalities with diverse granularity. Specifically, we devise a multi-grained video encoder that employs spatial, temporal, and global prompts to transfer the base-generic knowledge from the image-text pre-trained model while comprehensively excavating determinative video-specific characteristics. Meanwhile, we introduce a novel multi-grained text encoder aimed at capturing various levels of textual clues through the utilization of word and phrase prompts. Extensive experiments on four benchmark datasets, i.e., MSR-VTT, ActivityNet, DiDeMo, and LSMDC, demonstrate that MPT achieves outstanding performance, surpassing state-of-the-art methods with negligible computational cost. The codebase is publicly available at: https://github.com/zchoi/MPT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper9
- Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video RetrievalJian Xiao, Zijie Song, Jialong Hu, Hao Cheng 等NeurIPS 2025 · 被引用 3 次
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingJiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu 等NeurIPS 2025 · 被引用 2 次
- CARIM: Caption-Based Autonomous Driving Scene Retrieval via Inclusive Text MatchingMinjoo Ki, Daejung Kim, Kisung Kim, Seon Joo Kim 等ICCV 2025 · 被引用 2 次
- NeighborRetr: Balancing Hub Centrality in Cross-Modal RetrievalZengrong Lin, Zheng Wang, Tianwen Qian, Pan Mu 等CVPR 2025
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu 等SIGIR 2026
相关 Paper
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 被引用 53 次
- MV-Adapter: Multimodal Video Transfer Learning for Video Text RetrievalXiaojie Jin, Bowen Zhang, Weibo Gong, Kai Xu 等CVPR 2024
- VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal RetrievalSiteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang 等CVPR 2023
- PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalPeiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu 等ICCV 2023 · 被引用 25 次
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen 等ICCV 2023 · 被引用 52 次
