BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
Yulu Pan, Ce Zhang, Gedas Bertasius
Abstract
We present BASKET, a large-scale basketball video dataset for fine-grained skill estimation. BASKET contains 4,477 hours of video capturing 32,232 basketball players from all over the world. Compared to prior skill estimation datasets, our dataset includes a massive number of skilled participants with unprecedented diversity in terms of gender, age, skill level, geographical location, etc. BAS-KET includes 20 fine-grained basketball skills, challenging modern video recognition models to capture the intricate nuances of player skill through in-depth video analysis. Given a long highlight video (8-10 minutes) of a particular player, the model needs to predict the skill level (e.g., excellent, good, average, fair, poor) for each of the 20 basketball skills. Our empirical analysis reveals that the current state-of-the-art video models struggle with this task, significantly lagging behind the human baseline. We believe that BASKET could be a useful resource for developing new video models with advanced long-range, finegrained recognition capabilities. In addition, we hope that our dataset will be useful for domain-specific applications such as fair basketball scouting, personalized player development, and many others. Dataset and code are available at https://github.com/yulupan00/BASKET .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86b66c09-80eb-4dc0-b036-6ceeb221d6e9Cited by top-tier papers6
- SoccerMaster: A Vision Foundation Model for Soccer UnderstandingHaolin Yang, Jiayuan Rao, Haoning Wu, Weidi XieCVPR 2026 · 10 citations
- Learning Skill-Attributes for Transferable Assessment in VideoKumar Ashutosh, Kristen GraumanNeurIPS 2025 · 6 citations
- SkillSight: Efficient First-Person Skill Assessment with GazeChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 4 citations
- BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality AssessmentKanglei Zhou, Chang Li, Qingyi Pan, Liyuan WangCVPR 2026 · 3 citations
- VideoNet: A Large-Scale Dataset for Domain-Specific Action RecognitionTanush Yadav, Reza Salehi, Jae Sung Park, Vivek Ramanujan et al.CVPR 2026
Builds on24
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
Related papers
- FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action UnderstandingJinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou et al.CVPR 2024 · 12 citations
- SportsHHI: A Dataset for Human-Human Interaction Detection in Sports VideosTao Wu, Runyu He, Gangshan Wu, Limin WangCVPR 2024 · 9 citations
- A Descriptive Basketball Highlight Dataset for Automatic Commentary GenerationBenhui Zhang, Junyu Gao, Yuan YuanACM MM 2024 · 14 citations
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesKristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et al.CVPR 2024
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami et al.ICCV 2025 · 1 citation
