Lune

ACM MM2020Top-tier venue

Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal Contexts

Xun Yang, Xueliang Liu, Meng Jian, Xinjian Gao, Meng Wang

2020Year
47Citations
17Top-tier citations

Abstract

Grounding objects in visual context from natural language queries is a crucial yet challenging vision-and-language task, which has gained increasing attention in recent years. Existing work has primarily investigated this task in the context of still images. Despite their effectiveness, these methods cannot be directly migrated into the video context, mainly due to 1) the complex spatio-temporal structure of videos and 2) the scarcity of fine-grained annotations of videos. To effectively ground objects in videos is profoundly more challenging and less explored.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 14adb259-1dd4-443d-9a1b-ebaa2f9b3b6e

Cited by top-tier papers17

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines