Adaptively Building a Video-language Model for Video Captioning and Retrieval without Massive Video Pretraining
Zihao Liu, Xiaoyu Wu, Shengjin Wang, Jiayao Qian
Abstract
Large-scale pretrained image-language models have shown remarkable performance recently. However, building a video-language model is more challenging due to the complexity of video and the difficulty of collecting high-quality data. This paper builds a video-language model in an adaptive manner, which transfers the knowledge from the image domain and can achieve state-of-the-art performance without any further massive video pretraining. The main contributions include a Visual Perception Adapter that seamlessly and efficiently adapts a pretrained image-language model to the video domain and a fine-grained contrastive learning with Inter-modal Token Alignment that bridges semantic gaps between vision, audio, and language with less data. The proposed model is evaluated on video captioning and retrieval. Experiments demonstrate that the proposed model exhibits competitive performance compared to models pretrained on millions of video-text pairs. Notably, our model's CIDEr and R@1 scores on the MSR-VTT dataset exceed the existing state-of-the-art by 6.3% and 1.3%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8da642df-f619-4e49-9099-08697c166e38Related papers
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
- Tem-adapter: Adapting Image-Text Pretraining for Video Question AnswerGuangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang et al.ICCV 2023 · 32 citations
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 10 citations
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu et al.CVPR 2024
- CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language AlignmentHongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu et al.ICLR 2023 · 53 citations
