Dual-AI: Dual-path Actor Interaction Learning for Group Activity Recognition
Mingfei Han, David Junhao Zhang, Yali Wang, Rui Yan, Lina Yao, Xiaojun Chang, Yu Qiao
Abstract
Learning spatial-temporal relation among multiple actors is crucial for group activity recognition. Different group activities often show the diversified interactions between actors in the video. Hence, it is often difficult to model complex group activities from a single view of spatial-temporal actor evolution. To tackle this problem, we propose a distinct Dual-path Actor Interaction (Dual-AI) framework, which flexibly arranges spatial and temporal transformers in two complementary orders, enhancing actor relations by integrating merits from different spatio-temporal paths. Moreover, we introduce a novel Multi-scale Actor Contrastive Loss (MAC-Loss) between two interactive paths of Dual-AI. Via self-supervised actor consistency in both frame and video levels, MAC-Loss can effectively distinguish individual actor representations to reduce action confusion among different actors. Consequently, our Dual-AI can boost group activity recognition by fusing such discriminative features of different actors. To evaluate the proposed approach, we conduct extensive experiments on the widely used benchmarks, including Volleyball [21], Collective Activity [II], and NBA datasets [49]. The proposed Dual-AI achieves state-of-the-art performance on all these datasets. It is worth noting the proposed Dual-AI with 50% training data outperforms a number of recent approaches with 100% training data. This confirms the generalization power of Dual-AI for group activity recognition, even under the challenging scenarios of limited supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbbff3a1-ef50-4b7c-a31a-632f52c46ce2Cited by top-tier papers8
- Mask Propagation for Efficient Video Semantic SegmentationYuetian Weng, Mingfei Han, Haoyu He, Mingjie Li et al.NeurIPS 2023 · 36 citations
- Interaction-aware Joint Attention Estimation Using People AttributesChihiro Nakatani, Hiroaki Kawashima, Norimichi UkitaICCV 2023 · 9 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- Bi-Causal: Group Activity Recognition via Bidirectional CausalityYouliang Zhang, Wenxuan Liu, Danni Xu, Zhuo Zhou et al.CVPR 2024 · 6 citations
- Learning Group Activity Features Through Person Attribute PredictionChihiro Nakatani, Hiroaki Kawashima, Norimichi UkitaCVPR 2024 · 4 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 62 citations
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du et al.ACM MM 2022 · 21 citations
- VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity RecognitionZhuming Wang, Yihao Zheng, Jiarui Li, Yaofei Wu et al.ACM MM 2025 · 1 citation
- Self-Supervised Representation Learning for Skeleton-Based Group Activity RecognitionCunling Bian, Wei Feng, Song WangACM MM 2022 · 11 citations
