Actor-Transformers for Group Activity Recognition
Kirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, Cees G. M. Snoek
Abstract
This paper strives to recognize individual actions and group activities from videos. While existing solutions for this challenging problem explicitly model spatial and temporal relationships based on location of individual actors, we propose an actor-transformer model able to learn and selectively extract information relevant for group activity recognition. We feed the transformer with rich actorspecific static and dynamic representations expressed by features from a 2D pose network and 3D CNN, respectively. We empirically study different ways to combine these representations and show their complementary benefits. Experiments show what is important to transform and how it should be transformed. What is more, actor-transformers achieve state-of-the-art results on two publicly available benchmarks for group activity recognition, outperforming the previous best published results by a considerable margin. * This paper is the product of work during an internship at Sportlogiq.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b19047c3-ad03-4e41-9838-ef7870a714f7Cited by top-tier papers28
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang et al.ICCV 2021 · 149 citations
- Spatio-Temporal Dynamic Inference Network for Group Activity RecognitionHangjie Yuan, Dong Ni, Mang WangICCV 2021 · 113 citations
- Dual-AI: Dual-path Actor Interaction Learning for Group Activity RecognitionMingfei Han, David Junhao Zhang, Yali Wang, Rui Yan et al.CVPR 2022 · 80 citations
- Learning Visual Context for Group Activity RecognitionHangjie Yuan, Dong NiAAAI 2021 · 77 citations
Related papers
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 62 citations
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du et al.ACM MM 2022 · 21 citations
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo et al.ICCV 2023 · 38 citations
- Relational Self-Attention: What's Missing in Attention for Video UnderstandingManjin Kim, Heeseung Kwon, Chunyu Wang, Suha Kwak et al.NeurIPS 2021 · 40 citations
