Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment
Angchi Xu, Wei-Shi Zheng
摘要
Weakly-supervised action segmentation is a task of learning to partition a long video into several action segments, where training videos are only accompanied by transcripts (ordered list of actions). Most of existing methods need to infer pseudo segmentation for training by serial alignment between all frames and the transcript, which is time-consuming and hard to be parallelized while training. In this work, we aim to escape from this inefficient alignment with massive but redundant frames, and instead to directly localize a few action transitions for pseudo segmentation generation, where a transition refers to the change from an action segment to its next adjacent one in the transcript. As the true transitions are submerged in noisy boundaries due to intra-segment visual variation, we propose a novel Action-Transition-Aware Boundary Alignment (ATBA) framework to efficiently and effectively filter out noisy boundaries and detect transitions. In addition, to boost the semantic learning in the case that noise is inevitably present in the pseudo segmentation, we also introduce video-level losses to utilize the trusted video-level supervision. Extensive experiments show the effectiveness of our approach on both performance and training speed. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Code is available at https://github.com/iSEE-Laboratory/CVPR24_ATBA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Causal Temporal Representation Learning with Nonstationary Sparse TransitionXiangchen Song, Zijian Li, Guangyi Chen, Yujia Zheng 等NeurIPS 2024 · 被引用 18 次
- CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationElena Belén Bueno-Benito, Mariella DimiccoliICCV 2025 · 被引用 3 次
- Hierarchical Action Learning for Weakly-Supervised Action SegmentationJunxian Huang, Ruichu Cai, Juntao Fang, Hao Zhu 等CVPR 2026 · 被引用 1 次
- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosZijia Lu, A S. M. Iftekhar, Gaurav Mittal, Tianjian Meng 等CVPR 2025
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial SettingsJiajie Zhang, Sören Schwertfeger, Alexander KleinerCVPR 2026
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd 等CVPR 2022 · 被引用 275 次
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou 等NeurIPS 2021 · 被引用 252 次
相关 Paper
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 被引用 35 次
- Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2022 · 被引用 16 次
- Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyZijia Lu, Ehsan ElhamifarCVPR 2022 · 被引用 22 次
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 被引用 39 次
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 被引用 33 次
