BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online Policies
Thomas Verelst, Tinne Tuytelaars
Abstract
In this paper we propose BlockCopy, a scheme that accelerates pretrained frame-based CNNs to process video more efficiently, compared to standard frame-by-frame processing. To this end, a lightweight policy network determines important regions in an image, and operations are applied on selected regions only, using custom block-sparse convolutions. Features of non-selected regions are simply copied from the preceding frame, reducing the number of computations and latency. The execution policy is trained using reinforcement learning in an online fashion without requiring ground truth annotations. Our universal framework is demonstrated on dense prediction tasks such as pedestrian detection, instance segmentation and semantic segmentation, using both state of the art (Center and Scale Predictor, MGAN, SwiftNet) and standard baseline networks (Mask-RCNN, DeepLabV3+). BlockCopy achieves significant FLOPS savings and inference speedup with minimal impact on accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed2f5e17-812f-40b6-bc0c-80910d73d706Cited by top-tier papers3
- DSRC: Learning Density-Insensitive and Semantic-Aware Collaborative Representation Against CorruptionsJingyu Zhang, Yilei Wang, Lang Qian, Peng Sun et al.AAAI 2025 · 13 citations
- PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge DevicesQihua Zhou, Song Guo, Jun Pan, Jiacheng Liang et al.AAAI 2023
- Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosYubin Hu, Yuze He, Yanghao Li, Jisheng Li et al.CVPR 2023
Builds on5
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer et al.ICCV 2019 · 216 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Online Model Distillation for Efficient Video InferenceRavi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan et al.ICCV 2019 · 131 citations
- Generalizable Pedestrian Detection: The Elephant in the RoomIrtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram et al.CVPR 2021
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster InferenceThomas Verelst, Tinne TuytelaarsCVPR 2020
Related papers
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin et al.CVPR 2022 · 31 citations
- Fast Video Object Segmentation via Dynamic Targeting NetworkLu Zhang, Zhe Lin, Jianming Zhang, Huchuan Lu et al.ICCV 2019 · 59 citations
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song et al.ICCV 2021 · 117 citations
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin et al.CVPR 2022 · 27 citations
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
