Search-Map-Search: A Frame Selection Paradigm for Action Recognition
Mingjun Zhao, Yakun Yu, Xiaoli Wang, Lei Yang, Di Niu
Abstract
Despite the success of deep learning in video understanding tasks, processing every frame in a video is computationally expensive and often unnecessary in real-time applications. Frame selection aims to extract the most informative and representative frames to help a model better understand video content. Existing frame selection methods either individually sample frames based on per-frame importance prediction, without considering interaction among frames, or adopt reinforcement learning agents to find representative frames in succession, which are costly to train and may lead to potential stability issues. To overcome the limitations of existing methods, we propose a Search-Map-Search learning paradigm which combines the advantages of heuristic search and supervised learning to select the best combination of frames from a video as one entity. By combining search with learning, the proposed method can better capture frame interactions while incurring a low inference overhead. Specifically, we first propose a hierarchical search method conducted on each training video to search for the optimal combination of frames with the lowest error on the downstream task. A feature mapping function is then learned to map the frames of a video to the representation of its target optimal frame combination. During inference, another search is performed on an unseen video to select a combination of frames whose feature representation is close to the projected feature representation. Extensive experiments based on several action recognition benchmarks demonstrate that our frame selection method effectively improves performance of action recognition models, and significantly outperforms a number of competitive baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae58408f-ad68-4383-83fd-970661f41253Cited by top-tier papers2
- The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionOtto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens et al.CVPR 2025
- H-MoRe: Learning Human-centric Motion Representation for Action AnalysisZhanbo Huang, Xiaoming Liu, Yu KongCVPR 2025
Builds on9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
Related papers
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 6 citations
- An Efficient Framework for Dense Video CaptioningMaitreya Suin, A. N. RajagopalanAAAI 2020 · 49 citations
- Keyframe Extraction from Motion Capture Sequences with Graph based Deep Reinforcement LearningClinton Mo, Kun Hu, Shaohui Mei, Zebin Chen et al.ACM MM 2021 · 19 citations
- Visual Tracking via Hierarchical Deep Reinforcement LearningDawei Zhang, Zhonglong Zheng, Riheng Jia, Minglu LiAAAI 2021 · 29 citations
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij et al.CVPR 2021
