3D CNNs With Adaptive Temporal Feature Resolutions
Mohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi, Ehsan Adeli, Luc Van Gool, Jürgen Gall
Abstract
While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there is no setting that is optimal for all input clips. In this work, we therefore introduce a differentiable Similarity Guided Sampling (SGS) module, which can be plugged into any existing 3D CNN architecture. SGS empowers 3D CNNs by learning the similarity of temporal features and grouping similar features together. As a result, the temporal feature resolution is not anymore static but it varies for each input video clip. By integrating SGS as an additional layer within current 3D CNNs, we can convert them into much more efficient 3D CNNs with adaptive temporal feature resolutions (ATFR). Our evaluations show that the proposed module improves the state-of-the-art by reducing the computational cost (GFLOPs) by half while preserving or even improving the accuracy. We evaluate our module by adding it to multiple state-of-the-art 3D CNNs on various datasets such as Kinetics-600, Kinetics-400, Mini-Kinetics, Something-Something V2, UCF101, and HMDB51.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29cc9b4a-437e-4bda-b603-77e538d1c91bCited by top-tier papers5
- Efficient Video Action Detection with Token Dropout and Context RefinementLei Chen, Zhan Tong, Yibing Song, Gangshan Wu et al.ICCV 2023 · 31 citations
- Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art EvaluationMichal Balazia, Philipp Müller, Ákos Levente Tánczos, August von Liechtenstein et al.ACM MM 2022 · 28 citations
- Identity-aware Graph Memory Network for Action DetectionJingcheng Ni, Jie Qin, Di HuangACM MM 2021 · 8 citations
- Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferWenxuan Liu, Zhuo Zhou, Xuemei Jia, Siyuan Yang et al.AAAI 2026 · 1 citation
- Smooth Regularization for Efficient Video RecognitionGil Goldman, Raja Giryes, Mahadev SatyanarayananNeurIPS 2025 · 1 citation
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- SCSampler: Sampling Salient Clips From Video for Efficient Action RecognitionBruno Korbar, Du Tran, Lorenzo TorresaniICCV 2019 · 257 citations
- Multi-Agent Reinforcement Learning Based Frame Sampling for Effective Untrimmed Video RecognitionWenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen et al.ICCV 2019 · 135 citations
- DynamoNet: Dynamic Action and Motion NetworkAli Diba, Vivek Sharma, Luc Van Gool, Rainer StiefelhagenICCV 2019 · 123 citations
Related papers
- Gate-Shift Networks for Video Action RecognitionSwathikiran Sudhakaran, Sergio Escalera, Oswald LanzCVPR 2020
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang et al.AAAI 2020 · 267 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 2 citations
