Fast Object Detection in Compressed Video
Shiyao Wang, Hongchao Lu, Zhidong Deng
摘要
Object detection in videos has drawn increasing attention since it is more practical in real scenarios. Most of the deep learning methods use CNNs to process each decoded frame in a video stream individually. However, the free of charge yet valuable motion information already embedded in the video compression format is usually overlooked. In this paper, we propose a fast object detection method by taking advantage of this with a novel Motion aided Memory Network (MMNet). The MMNet has two major advantages: 1) It significantly accelerates the procedure of feature extraction for compressed videos. It only need to run a complete recognition network for I-frames, i.e. a few reference frames in a video, and it produces the features for the following P frames (predictive frames) with a light weight memory network, which runs fast; 2) Unlike existing methods that establish an additional network to model motion of frames, we take full advantage of both motion vectors and residual errors that are freely available in video streams. To our best knowledge, the MMNet is the first work that investigates a deep convolutional detector on compressed videos. Our method is evaluated on the large-scale Ima-geNet VID dataset, and the results show that it is 3× times faster than single image detector R-FCN and 10× times faster than high-performance detector MANet at a minor accuracy loss. * The work was done when Shiyao Wang was at Tsinghua University.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang 等SIGCOMM 2020 · 被引用 264 次
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao 等AAAI 2022 · 被引用 35 次
- CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video AnalyticsJinwoo Hwang, Minsu Kim, Daeun Kim, Seungho Nam 等USENIX ATC 2022 · 被引用 29 次
- Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionYuan Tian, Guo Lu, Guangtao Zhai, Zhiyong GaoICCV 2023 · 被引用 29 次
- Accelerating Video Object Segmentation with Compressed VideoKai Xu, Angela YaoCVPR 2022 · 被引用 24 次
相关 Paper
- MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector PriorShili Zhou, Xuhao Jiang, Weimin Tan, Ruian He 等ACM MM 2023 · 被引用 8 次
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang 等ICCV 2019 · 被引用 109 次
- End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionCongcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong 等CVPR 2022 · 被引用 18 次
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du 等AAAI 2020 · 被引用 40 次
- Motion Adaptive Pose Estimation from Compressed VideosZhipeng Fan, Jun Liu, Yao WangICCV 2021 · 被引用 24 次
