Patchwork: A Patch-Wise Attention Network for Efficient Object Detection and Segmentation in Video Streams
Yuning Chai
Abstract
Recent advances in single-frame object detection and segmentation techniques have motivated a wide range of works to extend these methods to process video streams. In this paper, we explore the idea of hard attention aimed for latency-sensitive applications. Instead of reasoning about every frame separately, our method selects and only processes a small sub-window of the frame. Our technique then makes predictions for the full frame based on the sub-windows from previous frames and the update from the current sub-window. The latency reduction by this hard attention mechanism comes at the cost of degraded accuracy. We made two contributions to address this. First, we propose a specialized memory cell that recovers lost context when processing sub-windows. Secondly, we adopt a Q-learning-based policy training strategy that enables our approach to intelligently select the sub-windows such that the staleness in the memory hurts the performance the least. Our experiments suggest that our approach reduces the latency by approximately four times without significantly sacrificing the accuracy on the ImageNet VID video object detection dataset and the DAVIS video object segmentation dataset. We further demonstrate that we can reinvest the saved computation into other parts of the network, and thus resulting in an accuracy increase at a comparable computational cost as the original system and beating other recently proposed state-of-the-art methods in the low latency range.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7866bc52-b74e-4610-abaa-ed0526b9cce3Cited by top-tier papers4
- FlexPatch: Fast and Accurate Object Detection for On-device High-Resolution Live Video AnalyticsKichang Yang, Juheon Yi, Kyungjin Lee, Youngki LeeINFOCOM 2022 · 43 citations
- Glimpse-Attend-and-Explore: Self-Attention for Active Visual ExplorationSoroush Seifi, Abhishek Jha, Tinne TuytelaarsICCV 2021 · 11 citations
- Learning To Recommend Frame for Interactive Video Object Segmentation in the WildZhaoyuan Yin, Jia Zheng, Weixin Luo, Shenhan Qian et al.CVPR 2021
- Boggart: Towards General-Purpose Acceleration of Retrospective Video AnalyticsNeil Agarwal, Ravi NetravaliNSDI 2023
Related papers
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 91 citations
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo et al.CVPR 2023
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Fast Object Detection in Compressed VideoShiyao Wang, Hongchao Lu, Zhidong DengICCV 2019 · 68 citations
