Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning
Wei Zhang, Chaoqun Wan, Tongliang Liu, Xinmei Tian, Xu Shen, Jieping Ye
Abstract
Extending large image-text pre-trained models (e.g., CLIP) for video understanding has made significant advancements. To enable the capability of CLIP to perceive dynamic information in videos, existing works are dedicated to equipping the visual encoder with various temporal modules. However, these methods exhibit “asymmetry” between the visual and textual sides, with neither temporal descriptions in input texts nor temporal modules in text encoder. This limitation hinders the potential of language supervision emphasized in CLIP, and restricts the learning of temporal features, as the text encoder has demonstrated limited proficiency in motion understanding. To address this issue, we propose leveraging “MoTion-Enhanced Descriptions” (MoTED) to facilitate the extraction of distinctive temporal features in videos. Specifically, we first generate discriminative motion-related descriptions via querying GPT-4 to compare easy-confusing action categories. Then, we incorporate both the visual and textual encoders with additional perception modules to process the video frames and generated descriptions, respectively. Finally, we adopt a contrastive loss to align the visual and textual motion features. Extensive experiments on five benchmarks show that MoTED surpasses state-of-the-art methods with convincing gaps, laying a solid foundation for empowering CLIP with strong temporal modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd3ff7f2-e4d2-4df5-bff6-47dd18907cb9Cited by top-tier papers5
- Storyboard-guided Alignment for Fine-grained Video Action RecognitionEnqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong et al.NeurIPS 2025 · 3 citations
- Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection Via Cross-Domain Retrieval AugmentationRongpei Hong, Jian Lang, Ting Zhong, Fan ZhouICCV 2025 · 3 citations
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 2 citations
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan et al.CVPR 2025
- BDC-CLIP: Brownian Distance Covariance for Adapting CLIP to Action RecognitionFei Long, Xiaoou Li, Jiaming Lv, Haoyuan Yang et al.ICML 2025
Builds on51
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalPeiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu et al.ICCV 2023 · 25 citations
- Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression RecognitionLiupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu et al.ACM MM 2024 · 4 citations
- Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship DetectionShuo Yang, Yongqi Wang, Xiaofeng Ji, Xinxiao WuAAAI 2024 · 4 citations
- HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language AlignmentRuijia Wu, Ping Chen, Fei Shen, Shaoan Zhao et al.AAAI 2026 · 1 citation
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 3 citations
