Leveraging Long-Range Temporal Relationships Between Proposals for Video Object Detection
Mykhailo Shvets, Wei Liu, Alexander C. Berg
Abstract
Single-frame object detectors perform well on videos sometimes, even without temporal context. However, challenges such as occlusion, motion blur, and rare poses of objects are hard to resolve without temporal awareness. Thus, there is a strong need to improve video object detection by considering long-range temporal dependencies. In this paper, we present a light-weight modification to a single-frame detector that accounts for arbitrary long dependencies in a video. It improves the accuracy of a single-frame detector significantly with negligible compute overhead. The key component of our approach is a novel temporal relation module, operating on object proposals, that learns the similarities between proposals from different frames and selects proposals from past and/or future to support current proposals. Our final “causal" model, without any offline post-processing steps, runs at a similar speed as a single-frame detector and achieves state-of-the-art video object detection on ImageNet VID dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0d79ae2-63ac-4735-bc7a-2fdacd02cbfbCited by top-tier papers28
- TF-Blender: Temporal Feature Blender for Video Object DetectionYiming Cui, Liqi Yan, Zhiwen Cao, Dongfang LiuICCV 2021 · 171 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- Temporal ROI Align for Video Object RecognitionTao Gong, Kai Chen, Xinjiang Wang, Qi Chu et al.AAAI 2021 · 108 citations
- End-to-End Video Object Detection with Spatial-Temporal TransformersLu He, Qianyu Zhou, Xiangtai Li, Li Niu et al.ACM MM 2021 · 106 citations
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu et al.NeurIPS 2022 · 91 citations
Related papers
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang et al.ACM MM 2020 · 37 citations
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao et al.AAAI 2022 · 35 citations
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong et al.CVPR 2020
