Lune

CVPR2020Top-tier venue

Spatio-Temporal Graph for Video Captioning With Knowledge Distillation

Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, Juan Carlos Niebles

2020Year
45Top-tier citations

Abstract

Figure 1: How to understand and describe a scene from video input? We argue that a detailed understanding of spatiotemporal object interaction is crucial for this task. In this paper, we propose a spatio-temporal graph model to explicitly capture such information for video captioning. Yellow boxes represent object proposals from Faster R-CNN [12]. Red arrows denote directed temporal edges (for clarity, only the most relevant ones are shown), while blue lines indicate undirected spatial connections. Video sample from MSVD [3] with the caption "A cat jumps into a box." Best viewed in color.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 41772bdd-9bfb-42c4-9ab1-25ea90b19639

Cited by top-tier papers45

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines