AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video Recognition
Yulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang, Zihang Lai, Victor Kulikov, Nikita Orlov, Humphrey Shi, Gao Huang
摘要
Recent works have shown that the computational efficiency of video recognition can be significantly improved by reducing the spatial redundancy. As a representative work, the adaptive focus method (AdaFocus) has achieved a favorable trade-off between accuracy and inference speed by dynamically identifying and attending to the informative regions in each video frame. However, AdaFocus requires a complicated three-stage training pipeline (involving reinforcement learning), leading to slow convergence and is unfriendly to practitioners. This work reformulates the training of AdaFocus as a simple one-stage algorithm by introducing a differentiable interpolation-based patch selection operation, enabling efficient end-to-end optimization. We further present an improved training scheme to address the issues introduced by the one-stage formulation, including the lack of supervision, input diversity and training stability. Moreover, a conditional-exit technique is proposed to perform temporal adaptive computation on top of AdaFocus without additional training. Extensive experiments on six benchmark datasets (i.e., ActivityNet, FCVID, Mini-Kinetics, Something-Something V1&V2, and Jester) demonstrate that our model significantly outperforms the original AdaFocus and other competitive baselines, while being considerably more simple and efficient to train. Code is available at https://github.com/ LeapLabTHU/AdaFocusV2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Efficient Spatially Sparse Inference for Conditional GANs and Diffusion ModelsMuyang Li, Ji Lin, Chenlin Meng, Stefano Ermon 等NeurIPS 2022 · 被引用 66 次
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song 等CVPR 2022 · 被引用 60 次
- Dynamic Perceiver for Efficient Visual RecognitionYizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang 等ICCV 2023 · 被引用 45 次
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 被引用 44 次
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan 等CVPR 2024 · 被引用 42 次
它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
相关 Paper
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song 等ICCV 2021 · 被引用 117 次
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri 等ICLR 2021 · 被引用 70 次
- FrameExit: Conditional Early Exiting for Efficient Video RecognitionAmir Ghodrati, Babak Ehteshami Bejnordi, Amirhossein HabibianCVPR 2021
- Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action RecognitionSyed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan 等ICCV 2023 · 被引用 39 次
- Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory RetrievalJing Zhang, Zhikai Li, Xuewen Liu, Qingyi GuICLR 2026 · 被引用 5 次
