ViSTec: Video Modeling for Sports Technique Recognition and Tactical Analysis
Yuchen He, Zeqing Yuan, Yihong Wu, Liqi Cheng, Dazhen Deng, Yingcai Wu
Abstract
The immense popularity of racket sports has fueled substantial demand in tactical analysis with broadcast videos. However, existing manual methods require laborious annotation, and recent attempts leveraging video perception models are limited to low-level annotations like ball trajectories, overlooking tactics that necessitate an understanding of stroke techniques. State-of-the-art action segmentation models also struggle with technique recognition due to frequent occlusions and motion-induced blurring in racket sports videos. To address these challenges, We propose ViSTec, a Video-based Sports Technique recognition model inspired by human cognition that synergizes sparse visual data with rich contextual insights. Our approach integrates a graph to explicitly model strategic knowledge in stroke sequences and enhance technique recognition with contextual inductive bias. A two-stage action perception model is jointly trained to align with the contextual knowledge in the graph. Experiments demonstrate that our method outperforms existing models by a significant margin. Case studies with experts from the Chinese national table tennis team validate our model's capacity to automate analysis for technical actions and tactical strategies. More details are available at: https://ViSTec2024.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 541a9dde-32be-4b09-856e-855721062294Cited by top-tier papers3
- CLOT: Closed Loop Optimal Transport for Unsupervised Action SegmentationElena Belén Bueno-Benito, Mariella DimiccoliICCV 2025 · 3 citations
- F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from VideosZhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou et al.ICLR 2025
- SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language ModelsHaotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy et al.ICLR 2025
Builds on12
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
- ShuttleSpace: Exploring and Analyzing Movement Trajectory in Immersive VisualizationShuainan Ye, Chen Zhu-Tian, Xiangtong Chu, Yifan Wang et al.IEEE VIS 2020 · 94 citations
- TIVEE: Visual Exploration and Explanation of Badminton Tactics in Immersive VisualizationsXiangtong Chu, Xiao Xie, Shuainan Ye, Haolin Lu et al.IEEE VIS 2021 · 68 citations
Related papers
- EventAnchor: Reducing Human Interactions in Event Annotation of Racket Sports VideosDazhen Deng, Jiang Wu, Jiachen Wang, Yihong Wu et al.CHI 2021 · 29 citations
- RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket AnalysisLinfeng Dong, Yuchen Yang, Hao Wu, Wei Wang et al.AAAI 2026 · 1 citation
- Augmenting Sports Videos with VisCommentatorChen Zhu-Tian, Shuainan Ye, Xiangtong Chu, Haijun Xia et al.IEEE VIS 2021 · 60 citations
- Where Will Players Move Next? Dynamic Graphs and Hierarchical Fusion for Movement Forecasting in BadmintonKai-Shiang Chang, Wei-Yao Wang, Wen-Chih PengAAAI 2023 · 24 citations
- VisMimic: Integrating Motion Chain in Feedback Video Generation for Motor CoachingLiqi Cheng, Xiao Xie, Yiwei Peng, Minghao Feng et al.UIST 2025 · 2 citations
