Efficient Action Recognition via Dynamic Knowledge Propagation
Hanul Kim, Mihir Jain, Jun-Tae Lee, Sungrack Yun, Fatih Porikli
Abstract
Efficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To this end, we employ two networks of different capabilities that operate in tandem to efficiently recognize actions. Given a video, the lighter network processes more frames while the heavier one only processes a few. In order to enable the effective interaction between the two, we propose dynamic knowledge propagation based on a cross-attention mechanism. This is the main component of our framework that is essentially a student-teacher architecture, but as the teacher model continues to interact with the student model during inference, we call it a dynamic student-teacher framework. Through extensive experiments, we demonstrate the effectiveness of each component of our framework. Our method outperforms competing state-of-the-art methods on two video datasets: ActivityNet-v1.3 and Mini-Kinetics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8aedf8b6-7240-44e6-a355-5c2845fd067eCited by top-tier papers8
- DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionThanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo et al.CVPR 2022 · 70 citations
- AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video RecognitionYulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang et al.CVPR 2022 · 52 citations
- Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing CorrectionAndrey V. SavchenkoICML 2023 · 43 citations
- SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-trainingYuanze Lin, Chen Wei, Huiyu Wang, Alan L. Yuille et al.ICCV 2023 · 18 citations
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo et al.NeurIPS 2022 · 17 citations
Builds on9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
Related papers
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao et al.AAAI 2024 · 9 citations
- Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video ModelsYing Peng, Hongsen Ye, Changxin Huang, Xiping Hu et al.AAAI 2026
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action RecognitionYue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri et al.ICLR 2021 · 70 citations
- FrameExit: Conditional Early Exiting for Efficient Video RecognitionAmir Ghodrati, Babak Ehteshami Bejnordi, Amirhossein HabibianCVPR 2021
- Knowledge Integration Networks for Action RecognitionShiwen Zhang, Sheng Guo, Limin Wang, Weilin Huang et al.AAAI 2020 · 20 citations
