Patchwork: A Patch-Wise Attention Network for Efficient Object Detection and Segmentation in Video Streams
Yuning Chai
摘要
Recent advances in single-frame object detection and segmentation techniques have motivated a wide range of works to extend these methods to process video streams. In this paper, we explore the idea of hard attention aimed for latency-sensitive applications. Instead of reasoning about every frame separately, our method selects and only processes a small sub-window of the frame. Our technique then makes predictions for the full frame based on the sub-windows from previous frames and the update from the current sub-window. The latency reduction by this hard attention mechanism comes at the cost of degraded accuracy. We made two contributions to address this. First, we propose a specialized memory cell that recovers lost context when processing sub-windows. Secondly, we adopt a Q-learning-based policy training strategy that enables our approach to intelligently select the sub-windows such that the staleness in the memory hurts the performance the least. Our experiments suggest that our approach reduces the latency by approximately four times without significantly sacrificing the accuracy on the ImageNet VID video object detection dataset and the DAVIS video object segmentation dataset. We further demonstrate that we can reinvest the saved computation into other parts of the network, and thus resulting in an accuracy increase at a comparable computational cost as the original system and beating other recently proposed state-of-the-art methods in the low latency range.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FlexPatch: Fast and Accurate Object Detection for On-device High-Resolution Live Video AnalyticsKichang Yang, Juheon Yi, Kyungjin Lee, Youngki LeeINFOCOM 2022 · 被引用 43 次
- Glimpse-Attend-and-Explore: Self-Attention for Active Visual ExplorationSoroush Seifi, Abhishek Jha, Tinne TuytelaarsICCV 2021 · 被引用 11 次
- Learning To Recommend Frame for Interactive Video Object Segmentation in the WildZhaoyuan Yin, Jia Zheng, Weixin Luo, Shenhan Qian 等CVPR 2021
- Boggart: Towards General-Purpose Acceleration of Retrospective Video AnalyticsNeil Agarwal, Ravi NetravaliNSDI 2023
相关 Paper
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang 等ICCV 2019 · 被引用 109 次
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 被引用 91 次
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo 等CVPR 2023
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Fast Object Detection in Compressed VideoShiyao Wang, Hongchao Lu, Zhidong DengICCV 2019 · 被引用 68 次
