Video Action Segmentation via Contextually Refined Temporal Keypoints
Borui Jiang, Yang Jin, Zhentao Tan, Yadong Mu
Abstract
Video action segmentation involves categorizing each frame or short snippet of an untrimmed video into predefined action categories. Despite notable advancements in recent years, a considerable number of current approaches still rely on frame-wise segmentation that tends to render fragmentary results. To address it, we present an innovative approach for video action segmentation, centered around contextually refined temporal keypoints. Initially, our method identifies a set of sparse, over-complete temporal keypoints through non-local visual cues, with each keypoint representing a potential action segment candidate. Subsequent enhancements to these initial keypoints are achieved through iterative refining and re-assembling operations. Driven by the notion that optimal temporal keypoints should collectively resemble the true ground-truth structurally, we introduce a module that conducts graph matching between the keypoint-derived graph and the reference graph constructed from accurate annotations. This module effectively learns structural features used to further refine the initial keypoints. Moreover, a set of predefined rules is applied to re-assemble all temporal keypoints. The unfiltered temporal keypoints, resulting from these operations, are harnessed to generate the final action segments. We extensively evaluate our method across three video benchmarks: 50salads, GTEA, and Breakfast. Our proposed approach consistently demonstrates substantial improvements over existing methods, establishing its superiority in video action segmentation. It achieves F 1@50 scores (one of the key performance metrics for this task) of 79.5%, 83.4%, and 60.5%, respectively, v.s. previous state-of-the-art 78.5%, 79.8% and 57.4%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Efficient Temporal Action Segmentation via Boundary-aware Query VotingPeiyao Wang, Yuewei Lin, Erik Blasch, Jie Wei et al.NeurIPS 2024 · 30 citations
- Learning Solution-Aware Transformers for Efficiently Solving Quadratic Assignment ProblemZhentao Tan, Yadong MuICML 2024 · 5 citations
Builds on13
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Learning Combinatorial Embedding Networks for Deep Graph MatchingRunzhong Wang, Junchi Yan, Xiaokang YangICCV 2019 · 268 citations
- Deep Graph Matching ConsensusMatthias Fey, Jan Eric Lenssen, Christopher Morris, Jonathan Masci et al.ICLR 2020 · 227 citations
Related papers
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 74 citations
- Iterative Contrast-Classify for Semi-supervised Temporal Action SegmentationDipika Singhania, Rahul Rahaman, Angela YaoAAAI 2022 · 35 citations
- Temporally-Weighted Hierarchical Clustering for Unsupervised Action SegmentationM. Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba et al.CVPR 2021
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
- Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsZheng Xing, Weibing ZhaoAAAI 2024 · 18 citations
