PixelSieve: Towards Efficient Activity Analysis From Compressed Video Streams
Yongchen Wang, Ying Wang, Huawei Li, Xiaowei Li
Abstract
Pixel-level data redundancy in video induces additional memory and computing overhead when neural networks are employed to mine spatiotemporal patterns, e.g. activity and event labels from video streams. This work proposes PixelSieve, to enable highly efficient CNN-based activity analysis directly from video data in compressed formats. Instead of recovering original RGB frames from compressed video, PixelSieve utilizes the built-in metadata in compressed video streams to distill only the critical pixels that render relevant spatiotemporal features, and then conducts efficient CNN inference with the condensed inputs. PixelSieve removes the overhead of video decoding and significantly improves the performance of CNN-based video analysis by 4.5x on average. I.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- An Efficient Deep Learning Accelerator for Compressed Video AnalysisYongchen Wang, Ying Wang, Huawei Li, Yinhe Han et al.DAC 2020 · 4 citations
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang et al.ACM MM 2021 · 3 citations
- MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded ConvolutionsMathias Parger, Chengcheng Tang, Thomas Neff, Christopher D. Twigg et al.ICCV 2023 · 11 citations
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksZeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev et al.CVPR 2025
