Learning Action-guided Spatio-temporal Transformer for Group Activity Recognition
Wei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du, Jian-Jun Qiao
Abstract
Learning spatial and temporal relations among people plays an important role in recognizing group activity. Recently, transformer-based methods have become popular solutions due to the proposal of self-attention mechanism. However, the person-level features are fed directly into the self-attention module without any refinement. Moreover, group activity in a clip often involves unbalanced spatio-temporal interactions, where only a few persons with special actions are critical to identifying different activities. It is difficult to learn the spatio-temporal interactions due to the lack of elaborately modeling the action dependencies among all people. In this paper, a novel Action-guided Spatio-Temporal transFormer (ASTFormer) is proposed to capture the interaction relations for group activity recognition by learning action-centric aggregation and modeling spatio-temporal action dependencies. Specifically, ASTFormer starts with assigning all persons in each frame to the latent actions, while an action-centric aggregation strategy is performed by weighting the sum of residuals for each latent action under the supervision of global action information. Then, a dual-branch transformer is proposed to refine the inter- and intra-frame action-level features, where two encoders with the self-attention mechanism are employed to select important tokens. Next, a semantic action graph is explicitly devised to model the dynamic action-wise dependencies. Finally, our model is capable of boosting group activity recognition by fusing these important cues, while only requiring video-level action labels. Extensive experiments on two popular benchmarks (Volleyball and Collective Activity) demonstrate the superior performance of our method in comparison with the state-of-the-art methods using only raw RGB frames as input.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b589796e-7e58-4fbd-a698-c114eca38333Related papers
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang et al.ICCV 2021 · 149 citations
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 62 citations
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
- ASTA-Net: Adaptive Spatio-Temporal Attention Network for Person Re-Identification in VideosXierong Zhu, Jiawei Liu, Haoze Wu, Meng Wang et al.ACM MM 2020 · 10 citations
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
