PixelSieve: Towards Efficient Activity Analysis From Compressed Video Streams
Yongchen Wang, Ying Wang, Huawei Li, Xiaowei Li
摘要
Pixel-level data redundancy in video induces additional memory and computing overhead when neural networks are employed to mine spatiotemporal patterns, e.g. activity and event labels from video streams. This work proposes PixelSieve, to enable highly efficient CNN-based activity analysis directly from video data in compressed formats. Instead of recovering original RGB frames from compressed video, PixelSieve utilizes the built-in metadata in compressed video streams to distill only the critical pixels that render relevant spatiotemporal features, and then conducts efficient CNN inference with the condensed inputs. PixelSieve removes the overhead of video decoding and significantly improves the performance of CNN-based video analysis by 4.5x on average. I.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- An Efficient Deep Learning Accelerator for Compressed Video AnalysisYongchen Wang, Ying Wang, Huawei Li, Yinhe Han 等DAC 2020 · 被引用 4 次
- Video-to-Image Casting: A Flatting Method for Video AnalysisXu Chen, Chenqiang Gao, Feng Yang, Xiaohan Wang 等ACM MM 2021 · 被引用 3 次
- MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded ConvolutionsMathias Parger, Chengcheng Tang, Thomas Neff, Christopher D. Twigg 等ICCV 2023 · 被引用 11 次
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksZeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev 等CVPR 2025
