F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
Zhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou, Yun Lin, Jin Song Dong
Abstract
Analyzing Fast, Frequent, and Fine-grained (F 3 ) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F 3 criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F 3 Set, a benchmark that consists of video datasets for precise F 3 event detection. Datasets in F 3 Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently F 3 Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F 3 Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F 3 ED, for F 3 event detections, achieving superior performance. The dataset, model, and benchmark code are available at https: //github.com/F3Set/F3Set .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46696231-2145-4c1e-a2fa-a5b499ab9506Cited by top-tier papers2
- SoccerMaster: A Vision Foundation Model for Soccer UnderstandingHaolin Yang, Jiayuan Rao, Haoning Wu, Weidi XieCVPR 2026 · 10 citations
- AdaSpot: Spend Resolution Where It Matters for Precise Event SpottingArtur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert ClapésCVPR 2026 · 2 citations
Builds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
Related papers
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsYixuan Li, Lei Chen, Runyu He, Zhenzhi Wang et al.ICCV 2021 · 131 citations
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami et al.ICCV 2025 · 1 citation
- FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action UnderstandingJinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou et al.CVPR 2024 · 12 citations
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Improving LLM Video Understanding with 16 Frames Per SecondYixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang et al.ICML 2025
