Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order Consistency
Zijia Lu, Ehsan Elhamifar
摘要
We address the problem of set-supervised action learning, whose goal is to learn an action segmentation model using weak supervision in the form of sets of actions occurring in training videos. Our key observation is that videos within the same task have similar ordering of actions, which can be leveraged for effective learning. Therefore, we propose an attention-based method with a new Pairwise Ordering Consistency (POC) loss that encourages that for each common action pair in two videos of the same task, the attentions of actions follow a similar ordering. Unlike existing sequence alignment methods, which misalign actions in videos with different orderings or cannot reliably separate more from less consistent orderings, our POC loss efficiently aligns videos with different action orders and is differentiable, which enables end-to-end training. In addition, it avoids the time-consuming pseudo-label generation of prior works. Our method efficiently learns the actions and their temporal locations, therefore, extends the existing attention-based action localization methods from learning one action per video to multiple actions using our POC loss along with video-level and frame-level losses. By experiments on three datasets, we demonstrate that our method significantly improves the state of the art. We also show that our method, with a small modification, can effectively address the transcript-supervised action learning task, where actions and their ordering are available during training. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Code available at https://github.com/ZijiaLewisLu/CVPR22-POC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 被引用 51 次
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 被引用 36 次
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 被引用 33 次
- Causal Temporal Representation Learning with Nonstationary Sparse TransitionXiangchen Song, Zijian Li, Guangyi Chen, Yujia Zheng 等NeurIPS 2024 · 被引用 18 次
- Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2022 · 被引用 16 次
它引用的顶会 Paper19
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 被引用 234 次
- Weakly-supervised Temporal Action Localization by Uncertainty ModelingPilhyeon Lee, Jinglu Wang, Yan Lu, Hyeran ByunAAAI 2021 · 被引用 141 次
- Unsupervised Procedure Learning via Joint Dynamic SummarizationEhsan Elhamifar, Zwe NaingICCV 2019 · 被引用 61 次
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 被引用 35 次
相关 Paper
- SCT: Set Constrained Temporal Transformer for Set Supervised Action SegmentationMohsen Fayyaz, Jürgen GallCVPR 2020
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Set-Constrained Viterbi for Set-Supervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2020
- Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary AlignmentAngchi Xu, Wei-Shi ZhengCVPR 2024 · 被引用 8 次
- Similar Modality Enhancement and Action Consistency Learning for Weakly Supervised Temporal Action LocalizationMaodong Li, Chao Zheng, Jian Wang, Bing LiAAAI 2025 · 被引用 2 次
