Learning to Predict Activity Progress by Self-Supervised Video Alignment
Gerard Donahue, Ehsan Elhamifar
摘要
In this paper, we tackle the problem of self-supervised video alignment and activity progress prediction using inthe-wild videos. Our proposed self-supervised representation learning method carefully addresses different action orderings, redundant actions, and background frames to generate improved video representations compared to previous methods. Our model generalizes temporal cycleconsistency learning to allow for more flexibility in determining cycle-consistent neighbors. More specifically, to handle repeated actions, we propose a multi-neighbor cycle consistency and a multi-cycle-back regression loss by finding multiple soft nearest neighbors using a Gaussian Mixture Model. To handle background and redundant frames, we introduce a context-dependent drop function in our framework, discouraging the alignment of droppable frames. On the other hand, to learn from videos of multiple activities jointly, we propose a multi-head crosstask network, allowing us to embed a video and estimate progress without knowing its activity label. Experiments on multiple datasets show that our method outperforms the state-of-theart for video alignment and progress prediction. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 被引用 33 次
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin 等ICCV 2025 · 被引用 13 次
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 被引用 6 次
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 被引用 3 次
它引用的顶会 Paper20
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Weakly-supervised Temporal Action Localization by Uncertainty ModelingPilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran ByunAAAI 2021 · 被引用 141 次
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 被引用 87 次
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg 等NeurIPS 2021 · 被引用 78 次
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 被引用 64 次
相关 Paper
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 被引用 35 次
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga 等NeurIPS 2020 · 被引用 59 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 被引用 23 次
- Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityHanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen 等AAAI 2022 · 被引用 35 次
