Skip-Convolutions for Efficient Video Processing
Amirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami Bejnordi
Abstract
We propose Skip-Convolutions to leverage the large amount of redundancies in video streams and save computations. Each video is represented as a series of changes across frames and network activations, denoted as residuals. We reformulate standard convolution to be efficiently computed on residual frames: each layer is coupled with a binary gate deciding whether a residual is important to the model prediction, e.g. foreground regions, or it can be safely skipped, e.g. background regions. These gates can either be implemented as an efficient network trained jointly with convolution kernels, or can simply skip the residuals based on their magnitude. Gating functions can also incorporate block-wise sparsity structures, as required for efficient implementation on hardware platforms. By replacing all convolutions with Skip-Convolutions in two state-ofthe-art architectures, namely EfficientDet and HRNet, we reduce their computational cost consistently by a factor of 3 ∼ 4× for two different tasks, without any accuracy drop. Extensive comparisons with existing model compression, as well as image and video efficiency methods demonstrate that Skip-Convolutions set a new state-of-the-art by effectively exploiting the temporal redundancies in videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7b98da7-4106-4e79-9875-d78beeac11e1Cited by top-tier papers18
- Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingYiding Wang, Decang Sun, Kai Chen, Fan Lai et al.EuroSys 2023 · 43 citations
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin et al.CVPR 2022 · 31 citations
- AccDecoder: Accelerated Decoding for Neural-enhanced Video AnalyticsTingting Yuan, Liang Mi, Weijun Wang, Haipeng Dai et al.INFOCOM 2023 · 25 citations
- Video Super-Resolution Transformer with Masked Inter&Intra-Frame AttentionXingyu Zhou, Leheng Zhang, Xiaorui Zhao, Keze Wang et al.CVPR 2024 · 20 citations
- INS-Conv: Incremental Sparse Convolution for Online 3D SegmentationLeyao Liu, Tian Zheng, Yun-Jou Lin, Kai Ni et al.CVPR 2022 · 19 citations
Builds on7
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Batch-shaping for learning conditional channel gated networksBabak Ehteshami Bejnordi, Tijmen Blankevoort, Max WellingICLR 2020 · 82 citations
- Dynamic Kernel Distillation for Efficient Pose Estimation in VideosXuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang et al.ICCV 2019 · 76 citations
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
- EfficientDet: Scalable and Efficient Object DetectionMingxing Tan, Ruoming Pang, Quoc V. LeCVPR 2020
Related papers
- FASTER Recurrent Networks for Efficient Video ClassificationLinchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang et al.AAAI 2020 · 60 citations
- Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy ReductionChaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat et al.ICCV 2023 · 4 citations
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksZeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev et al.CVPR 2025
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin et al.ICLR 2021 · 20 citations
- Stochastic Backpropagation: A Memory Efficient Strategy for Training Video ModelsFeng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen et al.CVPR 2022 · 11 citations
