Learning by Aligning Videos in Time
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed, Andrey Konin, M. Zeeshan Zia, Quoc-Huy Tran
Abstract
Embedding Video 1 Embedding Video 2 Encoder Figure 1 : We propose a self-supervised method to learn video representations by aligning videos in time, despite many differences between the videos such as appearance, motion, and viewpoint. We optimize the embedding space by using both the temporal alignment loss between the videos and the temporal regularization applied separately on each video. Our learned representations can be useful for many video-based temporal understanding tasks such as temporal video alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 64 citations
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin et al.CVPR 2022 · 52 citations
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 34 citations
- Weakly-Supervised Online Action Segmentation in Multi-View Instructional VideosReza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi et al.CVPR 2022 · 22 citations
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh et al.CVPR 2022 · 18 citations
Builds on3
- DynamoNet: Dynamic Action and Motion NetworkAli Diba, Vivek Sharma, Luc Van Gool, Rainer StiefelhagenICCV 2019 · 123 citations
- Predicting the Future: A Jointly Learnt Model for Action AnticipationHarshala Gammulle, Simon Denman, Sridha Sridharan, Clinton FookesICCV 2019 · 93 citations
- Few-Shot Video Classification via Temporal AlignmentKaidi Cao, Jingwei Ji, Zhangjie Cao, Chien-Yi Chang et al.CVPR 2020
Related papers
- Enhancing Self-supervised Video Representation Learning via Multi-level Feature OptimizationRui Qian, Yuxi Li, Huabin Liu, John See et al.ICCV 2021 · 43 citations
- Learning to Align Sequential Actions in the WildWeizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet et al.CVPR 2022
- Broaden Your Views for Self-Supervised Video LearningAdrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang et al.ICCV 2021 · 139 citations
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 2 citations
- Learning to Predict Activity Progress by Self-Supervised Video AlignmentGerard Donahue, Ehsan ElhamifarCVPR 2024
