Look More but Care Less in Video Recognition
Yitian Zhang, Yue Bai, Huan Wang, Yi Xu, Yun Fu
摘要
Existing action recognition methods typically sample a few frames to represent each video to avoid the enormous computation, which often limits the recognition performance. To tackle this problem, we propose Ample and Focal Network (AFNet), which is composed of two branches to utilize more frames but with less computation. Specifically, the Ample Branch takes all input frames to obtain abundant information with condensed computation and provides the guidance for Focal Branch by the proposed Navigation Module; the Focal Branch squeezes the temporal size to only focus on the salient frames at each convolution block; in the end, the results of two branches are adaptively fused to prevent the loss of information. With this design, we can introduce more frames to the network but cost less computation. Besides, we demonstrate AFNet can utilize fewer frames while achieving higher accuracy as the dynamic selection in intermediate features enforces implicit temporal modeling. Further, we show that our method can be extended to reduce spatial redundancy with even less cost. Extensive experiments on five datasets demonstrate the effectiveness and efficiency of our method. Our code is available at https://github.com/BeSpontaneous/AFNet-pytorch .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- View while Moving: Efficient Video Recognition in Long-untrimmed VideosYe Tian, Mengyu Yang, Lanshan Zhang, Zhizhen Zhang 等ACM MM 2023 · 被引用 11 次
- Frame Flexible NetworkYitian Zhang, Yue Bai, Chang Liu, Huan Wang 等CVPR 2023
- Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in SportsYi Xu, Yun FuICLR 2025
它引用的顶会 Paper14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 被引用 257 次
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song 等ICCV 2021 · 被引用 117 次
- Adaptive Trajectory Prediction via Transferable GNNYi Xu, Lichen Wang, Yizhou Wang, Yun FuCVPR 2022 · 被引用 85 次
相关 Paper
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri 等ICLR 2021 · 被引用 70 次
- Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action RecognitionHuabin Liu, Weixian Lv, John See, Weiyao LinACM MM 2022 · 被引用 11 次
- AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video RecognitionYulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang 等CVPR 2022 · 被引用 52 次
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See 等AAAI 2020 · 被引用 18 次
- Selective Feature Compression for Efficient Activity Recognition InferenceChunhui Liu, Xinyu Li, Hao Chen, Davide Modolo 等ICCV 2021 · 被引用 10 次
