Improving Action Segmentation via Graph-Based Temporal Reasoning
Yifei Huang, Yusuke Sugano, Yoichi Sato
Abstract
Temporal relations among multiple action segments play an important role in action segmentation especially when observations are limited (e.g., actions are occluded by other objects or happen outside a field of view). In this paper, we propose a network module called Graph-based Temporal Reasoning Module (GTRM) that can be built on top of existing action segmentation models to learn the relation of multiple action segments in various time spans. We model the relations by using two Graph Convolution Networks (GCNs) where each node represents an action segment. The two graphs have different edge properties to account for boundary regression and classification tasks, respectively. By applying graph convolution, we can update each node's representation based on its relation with neighboring nodes. The updated representation is then used for improved action segmentation. We evaluate our model on the challenging egocentric datasets namely EGTEA and EPIC-Kitchens, where actions may be partially observed due to the viewpoint restriction. The results show that our proposed GTRM outperforms state-of-the-art action segmentation models by a large margin. We also demonstrate the effectiveness of our model on two third-person video datasets, the 50Salads dataset and the Breakfast dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1192450-9d98-437b-9e2d-1922a5ba14b2Cited by top-tier papers39
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
- Generic Event Boundary Detection: A Benchmark for Event SegmentationMike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram et al.ICCV 2021 · 91 citations
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 74 citations
- ACGNet: Action Complement Graph Network for Weakly-Supervised Temporal Action LocalizationZichen Yang, Jie Qin, Di HuangAAAI 2022 · 72 citations
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang et al.ICCV 2023 · 72 citations
Builds on10
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall et al.ICCV 2019 · 294 citations
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis et al.ICCV 2019 · 201 citations
- Language-Conditioned Graph Networks for Relational ReasoningRonghang Hu, Anna Rohrbach, Trevor Darrell, Kate SaenkoICCV 2019 · 183 citations
Related papers
- Temporal Relational Modeling with Self-Supervision for Action SegmentationDong Wang, Di Hu, Xingjian Li, Dejing DouAAAI 2021 · 63 citations
- Graph-Based High-Order Relation Modeling for Long-Term Action RecognitionJiaming Zhou, Kun-Yu Lin, Haoxin Li, Wei-Shi ZhengCVPR 2021
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 11 citations
- Iterative Contrast-Classify for Semi-supervised Temporal Action SegmentationDipika Singhania, Rahul Rahaman, Angela YaoAAAI 2022 · 35 citations
- Multi-Modal Multi-Action Video RecognitionZhensheng Shi, Ju Liang, Qianqian Li, Haiyong Zheng et al.ICCV 2021 · 11 citations
