Temporal Relational Modeling with Self-Supervision for Action Segmentation
Dong Wang, Di Hu, Xingjian Li, Dejing Dou
摘要
Temporal relational modeling in video is essential for human action understanding, such as action recognition and action segmentation. Although Graph Convolution Networks (GCNs) have shown promising advantages in relation reasoning on many tasks, it is still a challenge to apply graph convolution networks on long video sequences effectively. The main reason is that large number of nodes (i.e., video frames) makes GCNs hard to capture and model temporal relations in videos. To tackle this problem, in this paper, we introduce an effective GCN module, Dilated Temporal Graph Reasoning Module (DTGRM), designed to model temporal relations and dependencies between video frames at various time spans. In particular, we capture and model temporal relations via constructing multi-level dilated temporal graphs where the nodes represent frames from different moments in video. Moreover, to enhance temporal reasoning ability of the proposed model, an auxiliary self-supervised task is proposed to encourage the dilated temporal graph reasoning module to find and correct wrong temporal relations in videos. Our DTGRM model outperforms state-of-the-art action segmentation models on three challenging datasets: 50Salads, Georgia Tech Egocentric Activities (GTEA), and the Breakfast dataset. The code is available at https://github.com/redwang/DTGRM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- How Much Temporal Long-Term Context is Needed for Action Segmentation?Emad Bahrami Rad, Gianpiero Francesca, Juergen GallICCV 2023 · 被引用 54 次
- Efficient Temporal Action Segmentation via Boundary-aware Query VotingPeiyao Wang, Yuewei Lin, Erik Blasch, Jie Wei 等NeurIPS 2024 · 被引用 30 次
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 被引用 5 次
- Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic ConditioningHao Zheng, Hu Wang, Tiantian Zheng, Prajjwal Bhattarai 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
- Graph-Based High-Order Relation Modeling for Long-Term Action RecognitionJiaming Zhou, Kun-Yu Lin, Haoxin Li, Wei-Shi ZhengCVPR 2021
- Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long VideosHaorui Wang, Yibo Hu, Yangfu Zhu, Jinsheng Qi 等ACM MM 2023 · 被引用 5 次
- Multi-Modal Multi-Action Video RecognitionZhensheng Shi, Ju Liang, Qianqian Li, Haiyong Zheng 等ICCV 2021 · 被引用 11 次
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 被引用 4 次
