Lune

ICCV2019Top-tier venue

Leveraging Long-Range Temporal Relationships Between Proposals for Video Object Detection

Mykhailo Shvets, Wei Liu, Alexander C. Berg

2019Year
91Citations
28Top-tier citations

Abstract

Single-frame object detectors perform well on videos sometimes, even without temporal context. However, challenges such as occlusion, motion blur, and rare poses of objects are hard to resolve without temporal awareness. Thus, there is a strong need to improve video object detection by considering long-range temporal dependencies. In this paper, we present a light-weight modification to a single-frame detector that accounts for arbitrary long dependencies in a video. It improves the accuracy of a single-frame detector significantly with negligible compute overhead. The key component of our approach is a novel temporal relation module, operating on object proposals, that learns the similarities between proposals from different frames and selects proposals from past and/or future to support current proposals. Our final “causal" model, without any offline post-processing steps, runs at a similar speed as a single-frame detector and achieves state-of-the-art video object detection on ImageNet VID dataset.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c0d79ae2-63ac-4735-bc7a-2fdacd02cbfb

Cited by top-tier papers28

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines