SMART Frame Selection for Action Recognition
Shreyank N. Gowda, Marcus Rohrbach, Laura Sevilla-Lara
Abstract
Action recognition is computationally expensive. In this paper, we address the problem of frame selection to improve the accuracy of action recognition. In particular, we show that selecting good frames helps in action recognition performance even in the trimmed videos domain. Recent work has successfully leveraged frame selection for long, untrimmed videos, where much of the content is not relevant, and easy to discard. In this work, however, we focus on the more standard short, trimmed action recognition problem. We argue that good frame selection can not only reduce the computational cost of action recognition but also increase the accuracy by getting rid of frames that are hard to classify. In contrast to previous work, we propose a method that instead of selecting frames by considering one at a time, considers them jointly. This results in a more efficient selection, where "good" frames are more effectively distributed over the video, like snapshots that tell a story. We call the proposed frame selection SMART and we test it in combination with different backbone architectures and on multiple benchmarks (Kinetics, Something-something, UCF101). We show that the SMART frame selection consistently improves the accuracy compared to other frame selection strategies while reducing the computational cost by a factor of 4 to 10 times. Additionally, we show that when the primary goal is recognition performance, our selection strategy can improve over recent state-of-the-art models and frame selection strategies on various benchmarks (UCF101, HMDB51, FCVID, and Activi-tyNet).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfcf6dfa-c22b-4354-b623-b9b868659f77Cited by top-tier papers14
- Context-Sensitive Temporal Feature Learning for Gait RecognitionXiaohu Huang, Duowang Zhu, Hao Wang, Xinggang Wang et al.ICCV 2021 · 159 citations
- Broaden Your Views for Self-Supervised Video LearningAdrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang et al.ICCV 2021 · 139 citations
- FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosYan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu et al.CVPR 2022 · 107 citations
- Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing CorrectionAndrey V. SavchenkoICML 2023 · 43 citations
- Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal PerceptionHassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung et al.NeurIPS 2023 · 33 citations
Builds on3
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
- Knowledge Integration Networks for Action RecognitionShiwen Zhang, Sheng Guo, Limin Wang, Weilin Huang et al.AAAI 2020 · 20 citations
Related papers
- Selective Feature Compression for Efficient Activity Recognition InferenceChunhui Liu, Xinyu Li, Hao Chen, Davide Modolo et al.ICCV 2021 · 10 citations
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij et al.CVPR 2021
- FrameExit: Conditional Early Exiting for Efficient Video RecognitionAmir Ghodrati, Babak Ehteshami Bejnordi, Amirhossein HabibianCVPR 2021
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin et al.CVPR 2022 · 27 citations
- 3D CNNs With Adaptive Temporal Feature ResolutionsMohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi et al.CVPR 2021
