Learning Action-guided Spatio-temporal Transformer for Group Activity Recognition
Wei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du, Jian-Jun Qiao
摘要
Learning spatial and temporal relations among people plays an important role in recognizing group activity. Recently, transformer-based methods have become popular solutions due to the proposal of self-attention mechanism. However, the person-level features are fed directly into the self-attention module without any refinement. Moreover, group activity in a clip often involves unbalanced spatio-temporal interactions, where only a few persons with special actions are critical to identifying different activities. It is difficult to learn the spatio-temporal interactions due to the lack of elaborately modeling the action dependencies among all people. In this paper, a novel Action-guided Spatio-Temporal transFormer (ASTFormer) is proposed to capture the interaction relations for group activity recognition by learning action-centric aggregation and modeling spatio-temporal action dependencies. Specifically, ASTFormer starts with assigning all persons in each frame to the latent actions, while an action-centric aggregation strategy is performed by weighting the sum of residuals for each latent action under the supervision of global action information. Then, a dual-branch transformer is proposed to refine the inter- and intra-frame action-level features, where two encoders with the self-attention mechanism are employed to select important tokens. Next, a semantic action graph is explicitly devised to model the dynamic action-wise dependencies. Finally, our model is capable of boosting group activity recognition by fusing these important cues, while only requiring video-level action labels. Extensive experiments on two popular benchmarks (Volleyball and Collective Activity) demonstrate the superior performance of our method in comparison with the state-of-the-art methods using only raw RGB frames as input.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang 等ICCV 2021 · 被引用 149 次
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 被引用 62 次
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang 等ACM MM 2020 · 被引用 10 次
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo 等ICCV 2023 · 被引用 38 次
