Identity-aware Graph Memory Network for Action Detection
Jingcheng Ni, Jie Qin, Di Huang
Abstract
Action detection plays an important role in high-level video understanding and media interpretation. Many existing studies fulfill this spatio-temporal localization by modeling the context, capturing the relationship of actors, objects, and scenes conveyed in the video. However, they often universally treat all the actors without considering the consistency and distinctness between individuals, leaving much room for improvement. In this paper, we explicitly highlight the identity information of the actors in terms of both long-term and short-term context through a graph memory network, namely identity-aware graph memory network (IGMN). Specifically, we propose the hierarchical graph neural network (HGNN) to comprehensively conduct long-term relation modeling within the same identity as well as between different ones. Regarding short-term context, we develop a dual attention module (DAM) to generate identity-aware constraint to reduce the influence of interference by the actors of different identities. Extensive experiments on the challenging AVA dataset demonstrate the effectiveness of our method, which achieves state-of-the-art results on AVA v2.1 and v2.2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58b6fd42-ca06-4bdd-9850-19b7441eb2d6Builds on6
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
- Actor-Context-Actor Relation Network for Spatio-Temporal Action LocalizationJunting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu et al.CVPR 2021
Related papers
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 57 citations
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- Graph Attention Based Proposal 3D ConvNets for Action DetectionJin Li, Xianglong Liu, Zhuofan Zong, Wanru Zhao et al.AAAI 2020 · 59 citations
- CAA: Candidate-Aware Aggregation for Temporal Action DetectionYifan Ren, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 3 citations
- Graph-Based High-Order Relation Modeling for Long-Term Action RecognitionJiaming Zhou, Kun-Yu Lin, Haoxin Li, Wei-Shi ZhengCVPR 2021
