SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
Haotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy, Yuqing Wang, Chi Lu, Christopher Lai, Yanjun He, Xun Shao, Zhuoqing Xie, Yuan-Fang Wang, Weining Shen, Hanjie Chen
Abstract
Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPORTU, a benchmark designed to assess MLLMs across multi-level sports reasoning tasks. SPORTU comprises two key components: SPORTU-text, featuring 900 multiple-choice questions with human-annotated explanations for rule comprehension and strategy understanding. This component focuses on testing models' ability to reason about sports solely through question-answering (QA), without requiring visual inputs; SPORTU-video, consisting of 1,701 slow-motion video clips across 7 different sports and 12,048 QA pairs, designed to assess multi-level reasoning, from simple sports recognition to complex tasks like foul detection and rule application. We evaluated four prevalent LLMs mainly utilizing few-shot learning paradigms supplemented by chain-of-thought (CoT) prompting on the SPORTU-text part. GPT-4o achieves the highest accuracy of 71%, but still falls short of human-level performance, highlighting room for improvement in rule comprehension and reasoning. The evaluation for the SPORTU-video part includes 6 proprietary and 8 open-source MLLMs. Experiments show that models fall short on hard tasks that require deep reasoning and rule-based understanding. GPT-4o performs the best with only 57.8% accuracy on the hard task, showing large room for improvement. We hope that SPORTU will serve as a critical step toward evaluating models' capabilities in sports understanding and reasoning. The dataset is available at https://github.com/chili-lab/SPORTU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46f8fc3a-b31c-4159-96d0-e88a9504a405Cited by top-tier papers9
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLMHanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang et al.ICLR 2026 · 64 citations
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang et al.CVPR 2026 · 12 citations
- SoccerMaster: A Vision Foundation Model for Soccer UnderstandingHaolin Yang, Jiayuan Rao, Haoning Wu, Weidi XieCVPR 2026 · 10 citations
- Multi-Agent System for Comprehensive Soccer UnderstandingJiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang et al.ACM MM 2025 · 5 citations
- VideoBrain: Learning Adaptive Frame Sampling for Long Video UnderstandingJunbo Zou, Ziheng Huang, Shengjie Zhang, Liwen Zhang et al.ICML 2026 · 4 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
Related papers
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in SportsHaotian Xia, Haonan Ge, Junbo Zou, Hyun Woo Choi et al.ICLR 2026 · 17 citations
- FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts ReasoningHaodong Chen, Haojian Huang, Xinxiang Yin, Dian ShaoACM MM 2025 · 1 citation
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language ModelsFanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu et al.ICLR 2025
- MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image ReasoningJiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang et al.ICLR 2026 · 7 citations
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in VideosXuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu et al.ICLR 2025
