MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions
Mathias Parger, Chengcheng Tang, Thomas Neff, Christopher D. Twigg, Cem Keskin, Robert Wang, Markus Steinberger
Abstract
Convolutional neural network inference on video input is computationally expensive and requires high memory bandwidth. Recently, DeltaCNN [26] managed to reduce the cost by only processing pixels with significant updates over the previous frame. However, DeltaCNN relies on static camera input. Moving cameras add new challenges in how to fuse newly unveiled image regions with already processed regions efficiently to minimize the update rate - without increasing memory overhead and without knowing the camera extrinsics of future frames. In this work, we propose MotionDeltaCNN, a sparse CNN inference framework that supports moving cameras. We introduce spherical buffers and padded convolutions to enable seamless fusion of newly unveiled regions and previously processed regions – without increasing memory footprint. Our evaluation shows that we outperform DeltaCNN by up to 90% for moving camera videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation ReuseJinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim et al.VLDB 2025 · 2 citations
- DeltaDEQ: Exploiting Heterogeneous Convergence for Accelerating Deep Equilibrium IterationsZuowen Wang, Longbiao Cheng, Pehuen Moure, Niklas Hahn et al.NeurIPS 2024 · 2 citations
Builds on2
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin et al.CVPR 2022 · 31 citations
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
Related papers
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksZeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev et al.CVPR 2025
- Ultrafast Video Attention Prediction with Coupled Knowledge DistillationKui Fu, Peipei Shi, Yafei Song, Shiming Ge et al.AAAI 2020 · 11 citations
- AdaMask: Enabling Machine-Centric Video Streaming with Adaptive Frame Masking for DNN Inference OffloadingShengzhong Liu, Tianshi Wang, Jinyang Li, Dachun Sun et al.ACM MM 2022 · 45 citations
- PixelSieve: Towards Efficient Activity Analysis From Compressed Video StreamsYongchen Wang, Ying Wang, Huawei Li, Xiaowei LiDAC 2021 · 1 citation
- MoViNets: Mobile Video Networks for Efficient Video RecognitionDan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang et al.CVPR 2021
