Lune

ACM MM2020顶会

Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal Contexts

Xun Yang, Xueliang Liu, Meng Jian, Xinjian Gao, Meng Wang

2020年份
47被引次数
17顶会引用

摘要

Grounding objects in visual context from natural language queries is a crucial yet challenging vision-and-language task, which has gained increasing attention in recent years. Existing work has primarily investigated this task in the context of still images. Despite their effectiveness, these methods cannot be directly migrated into the video context, mainly due to 1) the complex spatio-temporal structure of videos and 2) the scarcity of fine-grained annotations of videos. To effectively ground objects in videos is profoundly more challenging and less explored.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper17

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖