Compact Bilinear Augmented Query Structured Attention for Sport Highlights Classification
Yanbin Hao, Hao Zhang, Chong-Wah Ngo, Qiang Liu, Xiaojun Hu
Abstract
Understanding fine-grained activities, such as sport highlights, is a problem being overlooked and receives considerably less research attention. Potential reasons include absences of specific fine-grained action benchmark datasets, research preferences to general super-categorical activities classification, and challenges of large visual similarities between fine-grained actions. To tackle these, we collect and manually annotate two sport highlights datasets, i.e., Basketball-8 & Soccer-10, for fine-grained action classification. Sample clips in the datasets are annotated with professional sub-categorical actions like "dunk", "goalkeeping" and etc. We also propose a Compact Bilinear Augmented Query Structured Attention (CBA-QSA) module and stack it on top of general three-dimensional neural networks in a plug-and-play manner to emphasize important spatio-temporal clues in highlight clips. Specifically, we adapt the hierarchical attention neural networks, which contain learnable query-scheme, on the video to identify discriminative spatial/temporal visual clues within highlight clips. We name this altered attention which separately learns a query for spatial/temporal feature as query structured attention (QSA). Furthermore, we inflate bilinear mapping, which is a mature technique to represent local pairwise interactions for image-level fine-grained classification, on video understanding. In detail, we extend its compact version (i.e., compact bilinear mapping (CBM) based on TensorSketch) to deal with the three-dimensional video signal for modeling local pairwise motion information. We eventually incorporate CBM and QSA together to form CBA-QSA neural networks for fine-grained sport highlights classifications. Experimental results demonstrate that CBA-QSA improves the general state-of-the-arts on Basketball-8 and Soccer-10 datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 752ea98c-c280-48cc-a41d-0b4bd0535ce8Cited by top-tier papers6
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 48 citations
- Selective Dependency Aggregation for Action ClassificationYi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei et al.ACM MM 2021 · 31 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- Long-term Leap Attention, Short-term Periodic Shift for Video ClassificationHao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah NgoACM MM 2022 · 12 citations
- Hierarchical Hourglass Convolutional Network for Efficient Video ClassificationYi Tan, Yanbin Hao, Hao Zhang, Shuo Wang et al.ACM MM 2022 · 8 citations
Related papers
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
- Temporal Query Networks for Fine-Grained Video UnderstandingChuhan Zhang, Ankush Gupta, Andrew ZissermanCVPR 2021
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- Class Semantics-based Attention for Action DetectionDeepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li et al.ICCV 2021 · 77 citations
- Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long VideosLing-An Zeng, Fa-Ting Hong, Wei-Shi Zheng, Qi-Zhi Yu et al.ACM MM 2020 · 84 citations
