Search-Map-Search: A Frame Selection Paradigm for Action Recognition
Mingjun Zhao, Yakun Yu, Xiaoli Wang, Lei Yang, Di Niu
摘要
Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and representative frames to help a model better understand video content. Existing frame selection methods either individually sample frames based on per-frame importance prediction, without considering interaction among frames, or adopt reinforcement learning agents to find representative frames in succession, which are costly to train and may lead to potential stability issues. To overcome the limitations of existing methods, we propose a Search-Map-Search learning paradigm which combines the advantages of heuristic search and supervised learning to select the best combination of frames from a video as one entity. By combining search with learning, the proposed method can better capture frame interactions while incurring a low inference overhead. Specifically, we first propose a hierarchical search method conducted on each training video to search for the optimal combination of frames with the lowest error on the downstream task. A feature mapping function is then learned to map the frames of a video to the representation of its target optimal frame combination. During inference, another search is performed on an unseen video to select a combination of frames whose feature representation is close to the projected feature representation. Extensive experiments based on several action recognition benchmarks demonstrate that our frame selection method effectively improves performance of action recognition models, and significantly outperforms a number of competitive baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionOtto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens 等CVPR 2025
- H-MoRe: Learning Human-centric Motion Representation for Action AnalysisZhanbo Huang, Xiaoming Liu, Yu KongCVPR 2025
它引用的顶会 Paper9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
相关 Paper
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 被引用 6 次
- An Efficient Framework for Dense Video CaptioningMaitreya Suin, A. N. RajagopalanAAAI 2020 · 被引用 49 次
- Keyframe Extraction from Motion Capture Sequences with Graph based Deep Reinforcement LearningClinton Mo, Kun Hu, Shaohui Mei, Zebin Chen 等ACM MM 2021 · 被引用 19 次
- Visual Tracking via Hierarchical Deep Reinforcement LearningDawei Zhang, Zhonglong Zheng, Riheng Jia, Minglu LiAAAI 2021 · 被引用 29 次
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij 等CVPR 2021
