Learning Visual Context for Group Activity Recognition
Hangjie Yuan, Dong Ni
Abstract
Group activity recognition aims to recognize an overall activity in a multi-person scene. Previous methods strive to reason on individual features. However, they under-explore the person-specific contextual information, which is significant and informative in computer vision tasks. In this paper, we propose a new reasoning paradigm to incorporate global contextual information. Specifically, we propose two modules to bridge the gap between group activity and visual context. The first is Transformer based Context Encoding (TCE) module, which enhances individual representation by encoding global contextual information to individual features and refining the aggregated information. The second is Spatial-Temporal Bilinear Pooling (STBiP) module. It firstly further explores pairwise relationships for the context encoded individual representation, then generates semantic representations via gated message passing on a constructed spatial-temporal graph. On their basis, we further design a two-branch model that integrates the designed modules into a pipeline. Systematic experiments demonstrate each module's effectiveness on either branch. Visualizations indicate that visual contextual cues can be aggregated globally by TCE. Moreover, our method achieves state-of-the-art results on two widely used benchmarks using only RGB images as input and 2D backbones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 640f1794-fb7f-4e07-b0eb-58ab721b4d0bCited by top-tier papers13
- Spatio-Temporal Dynamic Inference Network for Group Activity RecognitionHangjie Yuan, Dong Ni, Mang WangICCV 2021 · 113 citations
- RLIP: Relational Language-Image Pre-training for Human-Object Interaction DetectionHangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng et al.NeurIPS 2022 · 88 citations
- Dual-AI: Dual-path Actor Interaction Learning for Group Activity RecognitionMingfei Han, David Junhao Zhang, Yali Wang, Rui Yan et al.CVPR 2022 · 80 citations
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie et al.ICCV 2023 · 69 citations
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 62 citations
Builds on4
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- GPS-Net: Graph Property Sensing Network for Scene Graph GenerationXin Lin, Changxing Ding, Jinquan Zeng, Dacheng TaoCVPR 2020
- Progressive Relation Learning for Group Activity RecognitionGuyue Hu, Bo Cui, Yuan He, Shan YuCVPR 2020
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
Related papers
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang et al.ICCV 2021 · 149 citations
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du et al.ACM MM 2022 · 21 citations
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 48 citations
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 6 citations
- Scene-Aware Context Reasoning for Unsupervised Abnormal Event Detection in VideosChe Sun, Yunde Jia, Yao Hu, Yuwei WuACM MM 2020 · 113 citations
