Learning to Predict Activity Progress by Self-Supervised Video Alignment
Gerard Donahue, Ehsan Elhamifar
Abstract
In this paper, we tackle the problem of self-supervised video alignment and activity progress prediction using inthe-wild videos. Our proposed self-supervised representation learning method carefully addresses different action orderings, redundant actions, and background frames to generate improved video representations compared to previous methods. Our model generalizes temporal cycleconsistency learning to allow for more flexibility in determining cycle-consistent neighbors. More specifically, to handle repeated actions, we propose a multi-neighbor cycle consistency and a multi-cycle-back regression loss by finding multiple soft nearest neighbors using a Gaussian Mixture Model. To handle background and redundant frames, we introduce a context-dependent drop function in our framework, discouraging the alignment of droppable frames. On the other hand, to learn from videos of multiple activities jointly, we propose a multi-head crosstask network, allowing us to embed a video and estimate progress without knowing its activity label. Experiments on multiple datasets show that our method outperforms the state-of-theart for video alignment and progress prediction. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9d6df81-e6ef-4401-8a71-6903f322f87cCited by top-tier papers10
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin et al.ICCV 2025 · 13 citations
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 6 citations
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
Builds on20
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Weakly-supervised Temporal Action Localization by Uncertainty ModelingPilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran ByunAAAI 2021 · 141 citations
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg et al.NeurIPS 2021 · 78 citations
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 64 citations
Related papers
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 35 citations
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga et al.NeurIPS 2020 · 59 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 23 citations
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen et al.AAAI 2022 · 35 citations
