TSA-Net: Tube Self-Attention Network for Action Quality Assessment
Shunli Wang, Dingkang Yang, Peng Zhai, Chixiao Chen, Lihua Zhang
Abstract
In recent years, assessing action quality from videos has attracted growing attention in computer vision community and humancomputer interaction. Most existing approaches usually tackle this problem by directly migrating the model from action recognition tasks, which ignores the intrinsic differences within the feature map such as foreground and background information. To address this issue, we propose a Tube Self-Attention Network (TSA-Net) for action quality assessment (AQA). Specifically, we introduce a single object tracker into AQA and propose the Tube Self-Attention Module (TSA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions. The TSA module is embedded in existing video networks to form TSA-Net. Overall, our TSA-Net is with the following merits: 1) High computational efficiency, 2) High flexibility, and 3) The state-of-theart performance. Extensive experiments are conducted on popular action quality assessment datasets including AQA-7 and MTL-AQA. Besides, a dataset named Fall Recognition in Figure Skating (FR-FS) is proposed to explore the basic action assessment in the figure skating scene. Our TSA-Net achieves the Spearman's Rank Correlation of 0.8476 and 0.9393 on AQA-7 and MTL-AQA, respectively, which are the new state-of-the-art results. The results on FR-FS also verify the effectiveness of the TSA-Net. The code and FR-FS dataset are publicly available at https:// github.com/ Shunli-Wang/ TSA-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4181fea-ce99-46e6-81b7-349825ac6538Cited by top-tier papers11
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionDingkang Yang, Shuai Huang, Zhi Xu, Zhenpeng Li et al.ICCV 2023 · 72 citations
- Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language SequencesDingkang Yang, Haopeng Kuang, Shuai Huang, Lihua ZhangACM MM 2022 · 64 citations
- Likert Scoring with Grade Decoupling for Long-term Action AssessmentAngchi Xu, Ling-An Zeng, Wei-Shi ZhengCVPR 2022 · 41 citations
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang et al.CVPR 2024 · 31 citations
Builds on7
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Action Assessment by Joint Relation GraphsJiahui Pan, Jibin Gao, Wei-Shi ZhengICCV 2019 · 141 citations
- Few-Shot Object Detection With Attention-RPN and Multi-Relation DetectorQi Fan, Wei Zhuo, Chi-Keung Tang, Yu-Wing TaiCVPR 2020
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
Related papers
- Hybrid Dynamic-static Context-aware Attention Network for Action Assessment in Long VideosLing-An Zeng, Fa-Ting Hong, Wei-Shi Zheng, Qi-Zhi Yu et al.ACM MM 2020 · 84 citations
- Group-aware Contrastive Regression for Action Quality AssessmentXumin Yu, Yongming Rao, Wenliang Zhao, Jiwen Lu et al.ICCV 2021 · 147 citations
- Uncertainty-Aware Score Distribution Learning for Action Quality AssessmentYansong Tang, Zanlin Ni, Jiahuan Zhou, Danyang Zhang et al.CVPR 2020
- Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentJianjun Xiang, Yuanjie Dang, Peng Chen, Ronghua Liang et al.ACM MM 2024
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
