Representation Learning via Global Temporal Alignment and Cycle-Consistency
Isma Hadji, Konstantinos G. Derpanis, Allan D. Jepson
摘要
We introduce a weakly supervised method for representation learning based on aligning temporal sequences (e.g., videos) of the same process (e.g., human action). The main idea is to use the global temporal ordering of latent correspondences across sequence pairs as a supervisory signal. In particular, we propose a loss based on scoring the optimal sequence alignment to train an embedding network. Our loss is based on a novel probabilistic path finding view of dynamic time warping (DTW) that contains the following three key features: (i) the local path routing decisions are contrastive and differentiable, (ii) pairwise distances are cast as probabilities that are contrastive as well, and (iii) our formulation naturally admits a global cycleconsistency loss that verifies correspondences. For evaluation, we consider the tasks of fine-grained action classification, few shot learning, and video synchronization. We report significant performance increases over previous methods. In addition, we report two applications of our temporal alignment framework, namely 3D pose reconstruction and fine-grained audio/visual retrieval.
capable of scaling up to large amounts of data yet supporting finer-grained video understanding.
As outlined in Fig. 1, given a set of paired sequences capturing the same process (e.g., tennis forehand) but highly varied (i.e., different participants, action executions, scenes, and camera viewpoints), our method trains an embedding network to support the recovery of their latent temporal alignment. We refer to our method as weakly supervised, as only readily available sequence-level labels (e.g., tennis forehand) are required to construct training pairings containing the same process. Given such a pair of sequences, we use their latent temporal alignment as a supervisory signal to learn fine-grained temporal distinctions.
Key to our proposed method is a novel dynamic time warping (DTW) formulation to score global alignments between paired sequences. DTW enforces a stronger constraint over simply considering local (soft) nearest neighbour correspondences [15], since the temporal ordering of the matches in the sequences are taken into account. We depart from previous differentiable DTW methods [30,8,6]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 被引用 87 次
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg 等NeurIPS 2021 · 被引用 78 次
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li 等ICCV 2023 · 被引用 77 次
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 被引用 64 次
- Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge AugmentationKun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas PadoyNeurIPS 2024 · 被引用 58 次
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- SpeedNet: Learning the Speediness in VideosSagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri 等CVPR 2020
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
相关 Paper
- Weakly-Supervised Temporal Action Alignment Driven by Unbalanced Spectral Fused Gromov-Wasserstein DistanceDixin Luo, Yutong Wang, Angxiao Yue, Hongteng XuACM MM 2022 · 被引用 8 次
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh 等CVPR 2022 · 被引用 18 次
- Learning Discriminative Prototypes With Dynamic Time WarpingXiaobin Chang, Frederick Tung, Greg MoriCVPR 2021
- TempCLR: Temporal Alignment Representation with Contrastive LearningYuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen 等ICLR 2023
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed 等CVPR 2021
