Skip-Convolutions for Efficient Video Processing
Amirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami Bejnordi
摘要
We propose Skip-Convolutions to leverage the large amount of redundancies in video streams and save computations. Each video is represented as a series of changes across frames and network activations, denoted as residuals. We reformulate standard convolution to be efficiently computed on residual frames: each layer is coupled with a binary gate deciding whether a residual is important to the model prediction, e.g. foreground regions, or it can be safely skipped, e.g. background regions. These gates can either be implemented as an efficient network trained jointly with convolution kernels, or can simply skip the residuals based on their magnitude. Gating functions can also incorporate block-wise sparsity structures, as required for efficient implementation on hardware platforms. By replacing all convolutions with Skip-Convolutions in two state-ofthe-art architectures, namely EfficientDet and HRNet, we reduce their computational cost consistently by a factor of 3 ∼ 4× for two different tasks, without any accuracy drop. Extensive comparisons with existing model compression, as well as image and video efficiency methods demonstrate that Skip-Convolutions set a new state-of-the-art by effectively exploiting the temporal redundancies in videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingYiding Wang, Decang Sun, Kai Chen, Fan Lai 等EuroSys 2023 · 被引用 43 次
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin 等CVPR 2022 · 被引用 31 次
- AccDecoder: Accelerated Decoding for Neural-enhanced Video AnalyticsTingting Yuan, Liang Mi, Weijun Wang, Haipeng Dai 等INFOCOM 2023 · 被引用 25 次
- Video Super-Resolution Transformer with Masked Inter&Intra-Frame AttentionXingyu Zhou, Leheng Zhang, Xiaorui Zhao, Keze Wang 等CVPR 2024 · 被引用 20 次
- INS-Conv: Incremental Sparse Convolution for Online 3D SegmentationLeyao Liu, Tian Zheng, Yun-Jou Lin, Kai Ni 等CVPR 2022 · 被引用 19 次
它引用的顶会 Paper7
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Batch-shaping for learning conditional channel gated networksBabak Ehteshami Bejnordi, Tijmen Blankevoort, Max WellingICLR 2020 · 被引用 82 次
- Dynamic Kernel Distillation for Efficient Pose Estimation in VideosXuecheng Nie, Yuncheng Li, Linjie Luo, Ning Zhang 等ICCV 2019 · 被引用 76 次
- X3D: Expanding Architectures for Efficient Video RecognitionChristoph FeichtenhoferCVPR 2020
- EfficientDet: Scalable and Efficient Object DetectionMingxing Tan, Ruoming Pang, Quoc V. LeCVPR 2020
相关 Paper
- FASTER Recurrent Networks for Efficient Video ClassificationLinchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang 等AAAI 2020 · 被引用 60 次
- Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy ReductionChaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat 等ICCV 2023 · 被引用 4 次
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksZeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev 等CVPR 2025
- VA-RED2: Video Adaptive Redundancy ReductionBowen Pan, Rameswar Panda, Camilo Luciano Fosco, Chung-Ching Lin 等ICLR 2021 · 被引用 20 次
- Stochastic Backpropagation: A Memory Efficient Strategy for Training Video ModelsFeng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen 等CVPR 2022 · 被引用 11 次
