Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
Christodoulos Kechris, Jonathan Dan, José Miranda, David Atienza
Abstract
Deep learning time-series processing often relies on convolutional neural networks with overlapping windows. This overlap allows the network to produce an output faster than the window length. However, it introduces additional computations. This work explores the potential to optimize computational efficiency during inference by exploiting convolution's shiftinvariance properties to skip the calculation of layer activations between successive overlapping windows. Although convolutions are shift-invariant, zero-padding and pooling operations, widely used in such networks, are not efficient and complicate efficient streaming inference. We introduce StreamiNNC, a strategy to deploy Convolutional Neural Networks for online streaming inference. We explore the adverse effects of zero padding and pooling on the accuracy of streaming inference, deriving theoretical error upper bounds for pooling during streaming. We address these limitations by proposing signal padding and pooling alignment and provide guidelines for designing and deploying models for StreamiNNC. We validate our method in simulated data and on three real-world biomedical signal processing applications. StreamiNNC achieves a low deviation between streaming output and normal inference for all three networks (2.03 -3.55% NRMSE). This work demonstrates that it is possible to linearly speed up the inference of streaming CNNs processing overlapping windows, negating the additional computation typically incurred by overlapping windows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b32c771-03ca-49a2-b6b4-d9cfcd11280aBuilds on7
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Fractional Skipping: Towards Finer-Grained Dynamic CNN InferenceJianghao Shen, Yue Wang, Pengfei Xu, Yonggan Fu et al.AAAI 2020 · 49 citations
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu et al.AAAI 2023 · 40 citations
- PENNI: Pruned Kernel Sharing for Efficient CNN InferenceShiyu Li, Edward Hanson, Hai Li, Yiran ChenICML 2020 · 23 citations
- MoViNets: Mobile Video Networks for Efficient Video RecognitionDan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang et al.CVPR 2021
Related papers
- Short-Term Memory ConvolutionsGrzegorz Stefanski, Krzysztof Arendt, Pawel Daniluk, Bartlomiej Jasik et al.ICLR 2023 · 67 citations
- Asynchronous Event Processing with Local-Shift Graph Convolutional NetworkLinhui Sun, Yifan Zhang, Jian Cheng, Hanqing LuAAAI 2023 · 2 citations
- Refining activation downsampling with SoftPoolAlexandros Stergiou, Ronald Poppe, Grigorios KalliatakisICCV 2021 · 195 citations
- zkCNN: Zero Knowledge Proofs for Convolutional Neural Network Predictions and AccuracyTianyi Liu, Xiang Xie, Yupeng ZhangCCS 2021 · 4 citations
- StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the MicrocontrollerHong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai YehNeurIPS 2023 · 17 citations
