TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen, Xudong Lin, Guangxing Han, Shih-Fu Chang
Abstract
Video representation learning has been successful in video-text pre-training for zero-shot transfer, where each sentence is trained to be close to the paired video clips in a common feature space. For long videos, given a paragraph of description where the sentences describe different segments of the video, by matching all sentence-clip pairs, the paragraph and the full video are aligned implicitly. However, such unit-level comparison may ignore global temporal context, which inevitably limits the generalization ability. In this paper, we propose a contrastive learning framework TempCLR to compare the full video and the paragraph explicitly. As the video/paragraph is formulated as a sequence of clips/sentences, under the constraint of their temporal order, we use dynamic time warping to compute the minimum cumulative cost over sentence-clip pairs as the sequence-level distance. To explore the temporal dynamics, we break the consistency of temporal succession by shuffling video clips w.r.t. temporal granularity. Then, we obtain the representations for clips/sentences, which perceive the temporal information and thus facilitate the sequence alignment. In addition to pre-training on the video and paragraph, our approach can also generalize on the matching between video instances. We evaluate our approach on video retrieval, action step localization, and few-shot action recognition, and achieve consistent performance gain over all three tasks. Detailed ablation studies are provided to justify the approach design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6ac7c28-c7d0-41ce-b828-47f08572349aCited by top-tier papers7
- Multi-granularity Correspondence Learning from Long-term Noisy VideosYijie Lin, Jie Zhang, Zhenyu Huang, Jia Liu et al.ICLR 2024 · 42 citations
- MoDE: CLIP Data Experts via ClusteringJiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li et al.CVPR 2024 · 8 citations
- Exploring Temporal Concurrency for Video-Language Representation LearningHeng Zhang, Daqing Liu, Zezhong Lv, Bing Su et al.ICCV 2023 · 6 citations
- F-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound ExaminationBin Pu, Xusheng Liang, Xinpeng Ding, Jinlin Wu et al.CVPR 2026
- Multimodal Causality-Driven Representation Learning for Generalizable Medical Image SegmentationXUSHENG LIANG, Lihua Zhou, Nianxin Li, miao xu et al.CVPR 2026
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 11 citations
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 34 citations
- Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosSixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo et al.CVPR 2023
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts LearningQiang Wang, Junlong Du, Ke Yan, Shouhong DingACM MM 2023 · 26 citations
