GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer
Shuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang, Shinan Liu, Jun Hou, Shuai Yi
Abstract
Group activity recognition is a crucial yet challenging problem, whose core lies in fully exploring spatial-temporal interactions among individuals and generating reasonable group representations. However, previous methods either model spatial and temporal information separately, or directly aggregate individual features to form group features. To address these issues, we propose a novel group activity recognition network termed GroupFormer. It captures spatial-temporal contextual information jointly to augment the individual and group representations effectively with a clustered spatial-temporal transformer. Specifically, our GroupFormer has three appealing advantages: (1) A tailor-modified Transformer, Clustered Spatial-Temporal Transformer, is proposed to enhance the individual representation and group representation. (2) It models the spatial and temporal dependencies integrally and utilizes decoders to build the bridge between the spatial and temporal information. (3) A clustered attention mechanism is utilized to dynamically divide individuals into multiple clusters for better learning activity-aware semantic representations. Moreover, experimental results show that the proposed framework outperforms state-of-the-art methods on the Volleyball dataset and Collective Activity dataset. Code is available at https://github.com/xueyee/GroupFormer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9bf8351-e476-4323-add3-c56508ebad4dCited by top-tier papers17
- Dual-AI: Dual-path Actor Interaction Learning for Group Activity RecognitionMingfei Han, David Junhao Zhang, Yali Wang, Rui Yan et al.CVPR 2022 · 80 citations
- RNTrajRec: Road Network Enhanced Trajectory Recovery with Spatial-Temporal TransformerYuqi Chen, Hanyuan Zhang, Weiwei Sun, Baihua ZhengICDE 2023 · 70 citations
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 62 citations
- JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity DetectionMahsa Ehsanpour, Fatemeh Sadat Saleh, Silvio Savarese, Ian D. Reid et al.CVPR 2022 · 49 citations
- Seeing the forest and the tree: Building representations of both individual and collective dynamics with transformersRan Liu, Mehdi Azabou, Max Dabagia, Jingyun Xiao et al.NeurIPS 2022 · 28 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Progressive Relation Learning for Group Activity RecognitionGuyue Hu, Bo Cui, Yuan He, Shan YuCVPR 2020
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
Related papers
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du et al.ACM MM 2022 · 21 citations
- GAFormer: Enhancing Timeseries Transformers Through Group-Aware EmbeddingsJingyun Xiao, Ran Liu, Eva L. DyerICLR 2024 · 13 citations
- Learning Visual Context for Group Activity RecognitionHangjie Yuan, Dong NiAAAI 2021 · 77 citations
- Self-Supervised Representation Learning for Skeleton-Based Group Activity RecognitionCunling Bian, Wei Feng, Song WangACM MM 2022 · 11 citations
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 6 citations
