BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online Policies
Thomas Verelst, Tinne Tuytelaars
摘要
In this paper we propose BlockCopy, a scheme that accelerates pretrained frame-based CNNs to process video more efficiently, compared to standard frame-by-frame processing. To this end, a lightweight policy network determines important regions in an image, and operations are applied on selected regions only, using custom block-sparse convolutions. Features of non-selected regions are simply copied from the preceding frame, reducing the number of computations and latency. The execution policy is trained using reinforcement learning in an online fashion without requiring ground truth annotations. Our universal framework is demonstrated on dense prediction tasks such as pedestrian detection, instance segmentation and semantic segmentation, using both state of the art (Center and Scale Predictor, MGAN, SwiftNet) and standard baseline networks (Mask-RCNN, DeepLabV3+). BlockCopy achieves significant FLOPS savings and inference speedup with minimal impact on accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DSRC: Learning Density-Insensitive and Semantic-Aware Collaborative Representation Against CorruptionsJingyu Zhang, Yilei Wang, Lang Qian, Peng Sun 等AAAI 2025 · 被引用 13 次
- PASS: Patch Automatic Skip Scheme for Efficient Real-Time Video Perception on Edge DevicesQihua Zhou, Song Guo, Jun Pan, Jiacheng Liang 等AAAI 2023
- Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosYubin Hu, Yuze He, Yanghao Li, Jisheng Li 等CVPR 2023
它引用的顶会 Paper5
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer 等ICCV 2019 · 被引用 216 次
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 被引用 170 次
- Online Model Distillation for Efficient Video InferenceRavi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan 等ICCV 2019 · 被引用 131 次
- Generalizable Pedestrian Detection: The Elephant in the RoomIrtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram 等CVPR 2021
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster InferenceThomas Verelst, Tinne TuytelaarsCVPR 2020
相关 Paper
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin 等CVPR 2022 · 被引用 31 次
- Fast Video Object Segmentation via Dynamic Targeting NetworkLu Zhang, Zhe Lin, Jianming Zhang, Huchuan Lu 等ICCV 2019 · 被引用 59 次
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song 等ICCV 2021 · 被引用 117 次
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 27 次
- Skip-Convolutions for Efficient Video ProcessingAmirhossein Habibian, Davide Abati, Taco S. Cohen, Babak Ehteshami BejnordiCVPR 2021
