F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
Zhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou, Yun Lin, Jin Song Dong
摘要
Analyzing Fast, Frequent, and Fine-grained (F 3 ) events presents a significant challenge in video analytics and multi-modal LLMs. Current methods struggle to identify events that satisfy all the F 3 criteria with high accuracy due to challenges such as motion blur and subtle visual discrepancies. To advance research in video understanding, we introduce F 3 Set, a benchmark that consists of video datasets for precise F 3 event detection. Datasets in F 3 Set are characterized by their extensive scale and comprehensive detail, usually encompassing over 1,000 event types with precise timestamps and supporting multi-level granularity. Currently F 3 Set contains several sports datasets, and this framework may be extended to other applications as well. We evaluated popular temporal action understanding methods on F 3 Set, revealing substantial challenges for existing techniques. Additionally, we propose a new method, F 3 ED, for F 3 event detections, achieving superior performance. The dataset, model, and benchmark code are available at https: //github.com/F3Set/F3Set .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SoccerMaster: A Vision Foundation Model for Soccer UnderstandingHaolin Yang, Jiayuan Rao, Haoning Wu, Weidi XieCVPR 2026 · 被引用 10 次
- AdaSpot: Spend Resolution Where It Matters for Precise Event SpottingArtur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert ClapésCVPR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
相关 Paper
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports ActionsYixuan Li, Lei Chen, Runyu He, Zhenzhi Wang 等ICCV 2021 · 被引用 131 次
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami 等ICCV 2025 · 被引用 1 次
- FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action UnderstandingJinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou 等CVPR 2024 · 被引用 12 次
- FineGym: A Hierarchical Video Dataset for Fine-Grained Action UnderstandingDian Shao, Yue Zhao, Bo Dai, Dahua LinCVPR 2020
- Improving LLM Video Understanding with 16 Frames Per SecondYixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang 等ICML 2025
