GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer
Shuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang, Shinan Liu, Jun Hou, Shuai Yi
摘要
Group activity recognition is a crucial yet challenging problem, whose core lies in fully exploring spatial-temporal interactions among individuals and generating reasonable group representations. However, previous methods either model spatial and temporal information separately, or directly aggregate individual features to form group features. To address these issues, we propose a novel group activity recognition network termed GroupFormer. It captures spatial-temporal contextual information jointly to augment the individual and group representations effectively with a clustered spatial-temporal transformer. Specifically, our GroupFormer has three appealing advantages: (1) A tailor-modified Transformer, Clustered Spatial-Temporal Transformer, is proposed to enhance the individual representation and group representation. (2) It models the spatial and temporal dependencies integrally and utilizes decoders to build the bridge between the spatial and temporal information. (3) A clustered attention mechanism is utilized to dynamically divide individuals into multiple clusters for better learning activity-aware semantic representations. Moreover, experimental results show that the proposed framework outperforms state-of-the-art methods on the Volleyball dataset and Collective Activity dataset. Code is available at https://github.com/xueyee/GroupFormer
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Dual-AI: Dual-path Actor Interaction Learning for Group Activity RecognitionMingfei Han, David Junhao Zhang, Yali Wang, Rui Yan 等CVPR 2022 · 被引用 80 次
- RNTrajRec: Road Network Enhanced Trajectory Recovery with Spatial-Temporal TransformerYuqi Chen, Hanyuan Zhang, Weiwei Sun, Baihua ZhengICDE 2023 · 被引用 70 次
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 被引用 62 次
- JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity DetectionMahsa Ehsanpour, Fatemeh Sadat Saleh, Silvio Savarese, Ian D. Reid 等CVPR 2022 · 被引用 49 次
- Seeing the forest and the tree: Building representations of both individual and collective dynamics with transformersRan Liu, Mehdi Azabou, Max Dabagia, Jingyun Xiao 等NeurIPS 2022 · 被引用 28 次
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- Progressive Relation Learning for Group Activity RecognitionGuyue Hu, Bo Cui, Yuan He, Shan YuCVPR 2020
- Actor-Transformers for Group Activity RecognitionKirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. SnoekCVPR 2020
相关 Paper
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du 等ACM MM 2022 · 被引用 21 次
- GAFormer: Enhancing Timeseries Transformers Through Group-Aware EmbeddingsJingyun Xiao, Ran Liu, Eva L. DyerICLR 2024 · 被引用 13 次
- Learning Visual Context for Group Activity RecognitionHangjie Yuan, Dong NiAAAI 2021 · 被引用 77 次
- Self-Supervised Representation Learning for Skeleton-Based Group Activity RecognitionCunling Bian, Wei Feng, Song WangACM MM 2022 · 被引用 11 次
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 被引用 6 次
