Refining Action Segmentation with Hierarchical Video Representations
Hyemin Ahn, Dongheui Lee
Abstract
In this paper, we propose Hierarchical Action Segmentation Refiner (HASR), which can refine temporal action segmentation results from various models by understanding the overall context of a given video in a hierarchical way. When a backbone model for action segmentation estimates how the given video can be segmented, our model extracts segment-level representations based on frame-level features, and extracts a video-level representation based on the segment-level representations. Based on these hierarchical representations, our model can refer to the overall context of the entire video, and predict how the segment labels that are out of context should be corrected. Our HASR can be plugged into various action segmentation models (MS-TCN, SSTDA, ASRF), and improve the performance of state-of-the-art models based on three challenging datasets (GTEA, 50Salads, and Breakfast). For example, in 50Salads dataset, the segmental edit score improves from 67.9% to 77.4% (MS-TCN), from 75.8% to 77.3% (SSTDA), from 79.3% to 81.0% (ASRF). In addition, our model can refine the segmentation result from the unseen backbone model, which was not referred to when training HASR. This generalization performance would make HASR be an effective tool for boosting up the existing approaches for temporal action segmentation. Our code is available at https: //github.com/cotton-ahn/HASR_iccv2021 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18b789da-e709-4813-99c7-b706e27a9d39Cited by top-tier papers20
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
- Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosMuheng Li, Lei Chen, Yueqi Duan, Zhilan Hu et al.CVPR 2022 · 70 citations
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Efficient Temporal Action Segmentation via Boundary-aware Query VotingPeiyao Wang, Yuewei Lin, Erik Blasch, Jie Wei et al.NeurIPS 2024 · 30 citations
- Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationZiwei Xu, Yogesh S. Rawat, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2022 · 18 citations
Builds on4
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Action Segmentation With Joint Self-Supervised Temporal Domain AdaptationMin-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib et al.CVPR 2020
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
Related papers
- Temporal Context Aggregation Network for Temporal Action Proposal RefinementZhiwu Qing, Haisheng Su, Weihao Gan, Dongliang Wang et al.CVPR 2021
- CAA: Candidate-Aware Aggregation for Temporal Action DetectionYifan Ren, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 3 citations
- Video Action Segmentation via Contextually Refined Temporal KeypointsBorui Jiang, Yang Jin, Zhentao Tan, Yadong MuICCV 2023 · 9 citations
- Progressive Boundary Refinement Network for Temporal Action DetectionQinying Liu, Zilei WangAAAI 2020 · 156 citations
- Temporally-Weighted Hierarchical Clustering for Unsupervised Action SegmentationM. Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba et al.CVPR 2021
