Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence Alignment
Yuhan Shen, Lu Wang, Ehsan Elhamifar
Abstract
We address the problem of unsupervised localization of task-relevant actions (key-steps) and feature learning in instructional videos using both visual and language instructions. Our key observation is that the sequences of visual and linguistic key-steps are weakly aligned: there is an ordered one-to-one correspondence between most visual and language key-steps, while some key-steps in one modality are absent in the other. To recover the two sequences, we develop an ordered prototype learning module, which extracts visual and linguistic prototypes representing key-steps. To find weak alignment and perform feature learning, we develop a differentiable weak sequence alignment (DWSA) method that finds ordered one-to-one matching between sequences while allowing some items in a sequence to stay unmatched. We develop an efficient forward and backward algorithm for computing the alignment and the loss derivative with respect to parameters of visual and language feature learning modules. By experiments on two instructional video datasets, we show that our method significantly improves the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 484edcf9-b743-42e6-9e0f-87608f88f016Cited by top-tier papers27
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 35 citations
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 28 citations
- Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyZijia Lu, Ehsan ElhamifarCVPR 2022 · 22 citations
Builds on11
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 109 citations
- Zero-Shot Anticipation for Instructional ActivitiesFadime Sener, Angela YaoICCV 2019 · 75 citations
- Unsupervised Procedure Learning via Joint Dynamic SummarizationEhsan Elhamifar, Zwe NaingICCV 2019 · 61 citations
Related papers
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg et al.NeurIPS 2021 · 78 citations
- StepFormer: Self-Supervised Step Discovery and Localization in Instructional VideosNikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G. Derpanis et al.CVPR 2023
- Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisTianyao He, Huabin Liu, Yuxi Li, Xiao Ma et al.AAAI 2024 · 8 citations
- Representation Learning via Global Temporal Alignment and Cycle-ConsistencyIsma Hadji, Konstantinos G. Derpanis, Allan D. JepsonCVPR 2021
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 87 citations
