SCT: Set Constrained Temporal Transformer for Set Supervised Action Segmentation
Mohsen Fayyaz, Jürgen Gall
Abstract
Temporal action segmentation is a topic of increasing interest, however, annotating each frame in a video is cumbersome and costly. Weakly supervised approaches therefore aim at learning temporal action segmentation from videos that are only weakly labeled. In this work, we assume that for each training video only the list of actions is given that occur in the video, but not when, how often, and in which order they occur. In order to address this task, we propose an approach that can be trained end-to-end on such data. The approach divides the video into smaller temporal regions and predicts for each region the action label and its length. In addition, the network estimates the action labels for each frame. By measuring how consistent the frame-wise predictions are with respect to the temporal regions and the annotated action labels, the network learns to divide a video into class-consistent regions. We evaluate our approach on three datasets where the approach achieves state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8817cda0-3beb-4cc9-b0f6-2ac7f4066505Cited by top-tier papers23
- MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly DetectionYingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton W. T. Fok et al.AAAI 2023 · 221 citations
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin et al.CVPR 2022 · 52 citations
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 39 citations
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 35 citations
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
Builds on3
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 109 citations
Related papers
- Temporal Action Segmentation From Timestamp SupervisionZhe Li, Yazan Abu Farha, Jürgen GallCVPR 2021
- Set-Constrained Viterbi for Set-Supervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2020
- PivoTAL: Prior-Driven Supervision for Weakly-Supervised Temporal Action LocalizationMamshad Nayeem Rizve, Gaurav Mittal, Ye Yu, Matthew Hall et al.CVPR 2023
- Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsYu Wang, Yadong Li, Hongbin WangCVPR 2023
- Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyZijia Lu, Ehsan ElhamifarCVPR 2022 · 22 citations
