Video Object Grounding Using Semantic Roles in Language Description
Arka Sadhu, Kan Chen, Ram Nevatia
Abstract
We explore the task of Video Object Grounding (VOG), which grounds objects in videos referred to in natural language descriptions. Previous methods apply image grounding based algorithms to address VOG, fail to explore the object relation information and suffer from limited generalization. Here, we investigate the role of object relations in VOG and propose a novel framework VOGNet to encode multi-modal object relations via self-attention with relative position encoding. To evaluate VOGNet, we propose novel contrasting sampling methods to generate more challenging grounding input samples, and construct a new dataset called ActivityNet-SRL (ASRL) based on existing caption and grounding datasets. Experiments on ASRL validate the need of encoding object relations in VOG, and our VOGNet outperforms competitive baselines by a significant margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 911c534f-a774-475c-a469-d7afb0d13121Cited by top-tier papers14
- SAT: 2D Semantics Assisted Training for 3D Visual GroundingZhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo LuoICCV 2021 · 166 citations
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 75 citations
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional VideosReuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin et al.NeurIPS 2021 · 30 citations
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
Builds on2
Related papers
- Video-Guided Curriculum Learning for Spoken Video GroundingYan Xia, Zhou Zhao, Shangwei Ye, Yang Zhao et al.ACM MM 2022 · 7 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan et al.CVPR 2025
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsJunli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang et al.ICCV 2025 · 5 citations
- Relation-aware Video Reading Comprehension for Temporal Language GroundingJialin Gao, Xin Sun, Mengmeng Xu, Xi Zhou et al.EMNLP 2021 · 51 citations
