Learning to Segment Actions from Observation and Narration
Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, Aida Nematzadeh
Abstract
We apply a generative segmental model of task structure, guided by narration, to action segmentation in video. We focus on unsupervised and weakly-supervised settings where no action labels are known during training. Despite its simplicity, our model performs competitively with previous work on a dataset of naturalistic instructional videos. Our model allows us to vary the sources of supervision used in training, and we find that both task structure and narrative language provide large benefits in segmentation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 28 citations
- Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyZijia Lu, Ehsan ElhamifarCVPR 2022 · 22 citations
- SVIP: Sequence VerIfication for Procedures in VideosYicheng Qian, Weixin Luo, Dongze Lian, Xu Tang et al.CVPR 2022 · 19 citations
- STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural VideosAnshul Shah, Benjamin Lundell, Harpreet Sawhney, Rama ChellappaICCV 2023 · 17 citations
Builds on2
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Unsupervised Procedure Learning via Joint Dynamic SummarizationEhsan Elhamifar, Zwe NaingICCV 2019 · 61 citations
Related papers
- Action Shuffle Alternating Learning for Unsupervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2021
- Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2022 · 16 citations
- SCT: Set Constrained Temporal Transformer for Set Supervised Action SegmentationMohsen Fayyaz, Jürgen GallCVPR 2020
- P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak SupervisionHe Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis et al.CVPR 2022 · 23 citations
- Temporally-Weighted Hierarchical Clustering for Unsupervised Action SegmentationM. Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba et al.CVPR 2021
