Representation Learning via Global Temporal Alignment and Cycle-Consistency
Isma Hadji, Konstantinos G. Derpanis, Allan D. Jepson
Abstract
We introduce a weakly supervised method for representation learning based on aligning temporal sequences (e.g., videos) of the same process (e.g., human action). The main idea is to use the global temporal ordering of latent correspondences across sequence pairs as a supervisory signal. In particular, we propose a loss based on scoring the optimal sequence alignment to train an embedding network. Our loss is based on a novel probabilistic path finding view of dynamic time warping (DTW) that contains the following three key features: (i) the local path routing decisions are contrastive and differentiable, (ii) pairwise distances are cast as probabilities that are contrastive as well, and (iii) our formulation naturally admits a global cycleconsistency loss that verifies correspondences. For evaluation, we consider the tasks of fine-grained action classification, few shot learning, and video synchronization. We report significant performance increases over previous methods. In addition, we report two applications of our temporal alignment framework, namely 3D pose reconstruction and fine-grained audio/visual retrieval.
capable of scaling up to large amounts of data yet supporting finer-grained video understanding.
As outlined in Fig. 1, given a set of paired sequences capturing the same process (e.g., tennis forehand) but highly varied (i.e., different participants, action executions, scenes, and camera viewpoints), our method trains an embedding network to support the recovery of their latent temporal alignment. We refer to our method as weakly supervised, as only readily available sequence-level labels (e.g., tennis forehand) are required to construct training pairings containing the same process. Given such a pair of sequences, we use their latent temporal alignment as a supervisory signal to learn fine-grained temporal distinctions.
Key to our proposed method is a novel dynamic time warping (DTW) formulation to score global alignments between paired sequences. DTW enforces a stronger constraint over simply considering local (soft) nearest neighbour correspondences [15], since the temporal ordering of the matches in the sequences are taken into account. We depart from previous differentiable DTW methods [30,8,6]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1152da8e-c458-4290-a049-42fb8df36e4cCited by top-tier papers22
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg et al.NeurIPS 2021 · 78 citations
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 64 citations
- Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge AugmentationKun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas PadoyNeurIPS 2024 · 58 citations
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- SpeedNet: Learning the Speediness in VideosSagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri et al.CVPR 2020
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
Related papers
- Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein DistanceDixin Luo, Yutong Wang, Angxiao Yue, Hongteng XuACM MM 2022 · 8 citations
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh et al.CVPR 2022 · 18 citations
- Learning Discriminative Prototypes With Dynamic Time WarpingXiaobin Chang, Frederick Tung, Greg MoriCVPR 2021
- TempCLR: Temporal Alignment Representation with Contrastive LearningYuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen et al.ICLR 2023
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed et al.CVPR 2021
