Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring
Ruyang Liu, Jingjia Huang, Ge Li, Jiashi Feng, Xinglong Wu, Thomas H. Li
Abstract
Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model, we revisit temporal modeling in the context of image-to-video knowledge transferring, which is the key point for extending image-text pretrained models to the video domain. We find that current temporal modeling mechanisms are tailored to either high-level semanticdominant tasks (e.g., retrieval) or low-level visual patterndominant tasks (e.g., recognition), and fail to work on the two cases simultaneously. The key difficulty lies in modeling temporal dependency while taking advantage of both highlevel and low-level knowledge in CLIP model. To tackle this problem, we present Spatial-Temporal Auxiliary Network (STAN) -a simple and effective temporal modeling mechanism extending CLIP model to diverse video tasks. Specifically, to realize both low-level and high-level knowledge transferring, STAN adopts a branch structure with decomposed spatial-temporal modules that enable multilevel CLIP features to be spatial-temporally contextualized. We evaluate our method on two representative video tasks: Video-Text Retrieval and Video Recognition. Extensive experiments demonstrate the superiority of our model over the state-of-the-art methods on various datasets, including MSR-VTT, DiDeMo, LSMDC, MSVD, Kinetics-400, and Something-Something-V2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- VideoPrism: A Foundational Visual Encoder for Video UnderstandingLong Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou et al.ICML 2024 · 91 citations
- FROSTER: Frozen CLIP is A Strong Teacher for Open-Vocabulary Action RecognitionXiaohu Huang, Hao Zhou, Kun Yao, Kai HanICLR 2024 · 56 citations
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with ToolsZhenlong Yuan, Xiangyan Qu, Chengxuan Qian, Rui Chen et al.ICLR 2026 · 32 citations
- PPLLaVA: Varied Video Sequence Understanding With Prompt GuidanceShangkun Sun, Ruyang Liu, Haoran Tang, Yixiao Ge et al.ICLR 2026 · 21 citations
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan et al.NeurIPS 2024 · 16 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language AlignmentHongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu et al.ICLR 2023 · 53 citations
- Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language ModelsWenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang et al.CVPR 2023
- T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video RetrievalYili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu et al.ACM MM 2025
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
- PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalPeiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu et al.ICCV 2023 · 25 citations
