AdaSpot: Spend Resolution Where It Matters for Precise Event Spotting
Artur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert Clapés
Abstract
Precise Event Spotting aims to localize fast-paced actions or events in videos with high temporal precision, a key task for applications in sports analytics, robotics, and autonomous systems. Existing methods typically process all frames uniformly, overlooking the inherent spatio-temporal redundancy in video data. This leads to redundant computation on non-informative regions while limiting overall efficiency. To remain tractable, they often spatially downsample inputs, losing fine-grained details crucial for precise localization. To address these limitations, we propose AdaSpot, a simple yet effective framework that processes low-resolution videos to extract global task-relevant features while adaptively selecting the most informative region-of-interest in each frame for high-resolution processing. The selection is performed via an unsupervised, task-aware strategy that maintains spatio-temporal consistency across frames and avoids the training instability of learnable alternatives. This design preserves essential fine-grained visual cues with a marginal computational overhead compared to low-resolution-only baselines, while remaining far more efficient than uniform high-resolution processing. Experiments on standard PES benchmarks demonstrate that AdaSpot achieves state-of-the-art performance under strict evaluation metrics (, and mAP frames on Tennis and FineDiving), while also maintaining strong results under looser metrics. Code is available at: https://github.com/arturxe2/AdaSpot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song et al.ICCV 2021 · 117 citations
Related papers
- Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalanceSanchayan Santra, Vishal M. Chudasama, Pankaj Wasnik, Vineeth N. BalasubramanianCVPR 2025
- A Context-Aware Loss Function for Action Spotting in Soccer VideosAnthony Cioppa, Adrien Deliège, Silvio Giancola, Bernard Ghanem et al.CVPR 2020
- Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and DistillationZhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou et al.AAAI 2026
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
- ExSample: Efficient Searches on Video Repositories through Adaptive SamplingOscar R. Moll, Favyen Bastani, Sam Madden, Mike Stonebraker et al.ICDE 2022 · 16 citations
