FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action Understanding
Jinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou, Yuxin Peng
摘要
Fine-grained action analysis in multi-person sports is complex due to athletes' quick movements and intense physical confrontations, which result in severe visual obstructions in most scenes. In addition, accessible multi-person sports video datasets lack fine- grained action annotations in both space and time, adding to the difficulty in fine- grained action analysis. To this end, we construct a new multi-person basketball sports video dataset named FineSports, which contains fine-grained semantic and spatial-temporal annotations on 10,000 NBA game videos, covering 52 fine-grained action types, 16,000 action instances, and 123,000 spatial-temporal bounding boxes. We also propose a new prompt-driven spatial-temporal action location approach called PoSTAL, composed of a prompt-driven target action encoder (PTA) and an action tube-specific detector (ATD) to directly generate target action tubes with fine-grained action types without any off-line proposal generation. Extensive experiments on the FineSports dataset demonstrate that PoSTAL outperforms state-of-the-art methods. Data and code are available at https://github.com/PKU-ICST-MIPL/FineSports_CVPR2024.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language ModelsHaotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy 等ICLR 2025
- VideoNet: A Large-Scale Dataset for Domain-Specific Action RecognitionTanush Yadav, Reza Salehi, Jae Sung Park, Vivek Ramanujan 等CVPR 2026
- Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Song Guo, Dacheng TaoCVPR 2025
- BASKET: A Large-Scale Video Dataset for Fine-Grained Skill EstimationYulu Pan, Ce Zhang, Gedas BertasiusCVPR 2025
它引用的顶会 Paper17
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
- SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports ScenesYutao Cui, Chenkai Zeng, Xiaoyu Zhao, Yichun Yang 等ICCV 2023 · 被引用 187 次
相关 Paper
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsYixuan Li, Lei Chen, Runyu He, Zhenzhi Wang 等ICCV 2021 · 被引用 131 次
- SportsHHI: A Dataset for Human-Human Interaction Detection in Sports VideosTao Wu, Runyu He, Gangshan Wu, Limin WangCVPR 2024 · 被引用 9 次
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen 等CVPR 2022 · 被引用 118 次
- F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from VideosZhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou 等ICLR 2025
- A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationBenhui Zhang, Junyu Gao, Yuan YuanACM MM 2024 · 被引用 14 次
