Look More but Care Less in Video Recognition
Yitian Zhang, Yue Bai, Huan Wang, Yi Xu, Yun Fu
Abstract
Existing action recognition methods typically sample a few frames to represent each video to avoid the enormous computation, which often limits the recognition performance. To tackle this problem, we propose Ample and Focal Network (AFNet), which is composed of two branches to utilize more frames but with less computation. Specifically, the Ample Branch takes all input frames to obtain abundant information with condensed computation and provides the guidance for Focal Branch by the proposed Navigation Module; the Focal Branch squeezes the temporal size to only focus on the salient frames at each convolution block; in the end, the results of two branches are adaptively fused to prevent the loss of information. With this design, we can introduce more frames to the network but cost less computation. Besides, we demonstrate AFNet can utilize fewer frames while achieving higher accuracy as the dynamic selection in intermediate features enforces implicit temporal modeling. Further, we show that our method can be extended to reduce spatial redundancy with even less cost. Extensive experiments on five datasets demonstrate the effectiveness and efficiency of our method. Our code is available at https://github.com/BeSpontaneous/AFNet-pytorch .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 143940e1-8b0b-4e5d-b620-dea00c714c4dCited by top-tier papers3
- View while Moving: Efficient Video Recognition in Long-untrimmed VideosYe Tian, Mengyu Yang, Lanshan Zhang, Zhizhen Zhang et al.ACM MM 2023 · 11 citations
- Frame Flexible NetworkYitian Zhang, Yue Bai, Chang Liu, Huan Wang et al.CVPR 2023
- Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in SportsYi Xu, Yun FuICLR 2025
Builds on14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song et al.ICCV 2021 · 117 citations
- Adaptive Trajectory Prediction via Transferable GNNYi Xu, Lichen Wang, Yizhou Wang, Yun FuCVPR 2022 · 85 citations
Related papers
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri et al.ICLR 2021 · 70 citations
- Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action RecognitionHuabin Liu, Weixian Lv, John See, Weiyao LinACM MM 2022 · 11 citations
- AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video RecognitionYulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang et al.CVPR 2022 · 52 citations
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
- Selective Feature Compression for Efficient Activity Recognition InferenceChunhui Liu, Xinyu Li, Hao Chen, Davide Modolo et al.ICCV 2021 · 10 citations
