Towards Universal Soccer Video Understanding
Jiayuan Rao, Haoning Wu, Hao Jiang, Ya Zhang, Yanfeng Wang, Weidi Xie
摘要
As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we make the following contributions in this paper: (i) we introduce SoccerReplay-1988, the largest multi-modal soccer dataset to date, featuring videos and detailed annotations from 1,988 complete matches, with an automated annotation pipeline; (ii) we present an advanced soccer-specific visual encoder, MatchVision, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks; (iii) we conduct extensive experiments and ablation studies on event classification, commentary generation, and multi-view foul recognition. MatchVision demonstrates state-of-the-art performance on all of them, substantially outperforming existing models, which highlights the superiority of our proposed data and model. We believe that this work will offer a standard paradigm for sports understanding research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Accelerating Streaming Video Large Language Models via Hierarchical Token CompressionYiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin 等CVPR 2026 · 被引用 40 次
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang 等CVPR 2026 · 被引用 12 次
- SoccerMaster: A Vision Foundation Model for Soccer UnderstandingHaolin Yang, Jiayuan Rao, Haoning Wu, Weidi XieCVPR 2026 · 被引用 10 次
- One Trajectory, One Token: Grounded Video Tokenization Via Panoptic Sub-Object TrajectoryChenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao 等ICCV 2025 · 被引用 7 次
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language ModelsJialiang Zhang, Junlong Tong, Junyan Lin, Hao Wu 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- MatchTime: Towards Automatic Soccer Game Commentary GenerationJiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang 等EMNLP 2024 · 被引用 14 次
- TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary GenerationLing You, Wenxuan Huang, Xinni Xie, Xiangyi Wei 等ACM MM 2025 · 被引用 2 次
- Multi-Agent System for Comprehensive Soccer UnderstandingJiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang 等ACM MM 2025 · 被引用 5 次
- Visual and Memory-Augmented Soccer Commentary GenerationHaoran Sun, Natthawut Kertkeidkachorn, Kiyoaki ShiraiACL 2026
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic 等EMNLP 2022 · 被引用 2 次
