Fast Object Detection in Compressed Video
Shiyao Wang, Hongchao Lu, Zhidong Deng
Abstract
Object detection in videos has drawn increasing attention since it is more practical in real scenarios. Most of the deep learning methods use CNNs to process each decoded frame in a video stream individually. However, the free of charge yet valuable motion information already embedded in the video compression format is usually overlooked. In this paper, we propose a fast object detection method by taking advantage of this with a novel Motion aided Memory Network (MMNet). The MMNet has two major advantages: 1) It significantly accelerates the procedure of feature extraction for compressed videos. It only need to run a complete recognition network for I-frames, i.e. a few reference frames in a video, and it produces the features for the following P frames (predictive frames) with a light weight memory network, which runs fast; 2) Unlike existing methods that establish an additional network to model motion of frames, we take full advantage of both motion vectors and residual errors that are freely available in video streams. To our best knowledge, the MMNet is the first work that investigates a deep convolutional detector on compressed videos. Our method is evaluated on the large-scale Ima-geNet VID dataset, and the results show that it is 3× times faster than single image detector R-FCN and 10× times faster than high-performance detector MANet at a minor accuracy loss. * The work was done when Shiyao Wang was at Tsinghua University.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb02f288-6a38-4eef-86a4-d0c23a53766aCited by top-tier papers17
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang et al.SIGCOMM 2020 · 264 citations
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao et al.AAAI 2022 · 35 citations
- CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video AnalyticsJinwoo Hwang, Minsu Kim, Daeun Kim, Seungho Nam et al.USENIX ATC 2022 · 29 citations
- Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic CompressionYuan Tian, Guo Lu, Guangtao Zhai, Zhiyong GaoICCV 2023 · 29 citations
- Accelerating Video Object Segmentation with Compressed VideoKai Xu, Angela YaoCVPR 2022 · 24 citations
Related papers
- MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector PriorShili Zhou, Xuhao Jiang, Weimin Tan, Ruian He et al.ACM MM 2023 · 8 citations
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionCongcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong et al.CVPR 2022 · 18 citations
- Temporal Context Enhanced Feature Aggregation for Video Object DetectionFei He, Naiyu Gao, Qiaozhe Li, Senyao Du et al.AAAI 2020 · 40 citations
- Motion Adaptive Pose Estimation from Compressed VideosZhipeng Fan, Jun Liu, Yao WangICCV 2021 · 24 citations
