Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in Videos
Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo, Shih-Fu Chang
Abstract
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking, proposal-based matching), we tackle the problem from a novel perspective, co-grounding, with an elegant one-stage framework. We enhance the single-frame grounding accuracy by semantic attention learning and improve the cross-frame grounding consistency with co-grounding feature learning. Semantic attention learning explicitly parses referring cues in different attributes to reduce the ambiguity in the complex expression. Co-grounding feature learning boosts visual feature representations by integrating temporal correlation to reduce the ambiguity caused by scene dynamics. Experiment results demonstrate the superiority of our framework on the video grounding datasets VID and LiOTB in generating accurate and stable results across frames. Our model is also applicable to referring expression comprehension in images, illustrated by the improved performance on the RefCOCO dataset. Our project is available at https://sijiesong.github.io/co- grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6eabaa1b-2d6d-4d18-87b8-9258ccf277aaCited by top-tier papers8
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Described Object Detection: Liberating Object Detection with Flexible ExpressionsChi Xie, Zhao Zhang, Yixuan Wu, Feng Zhu et al.NeurIPS 2023 · 69 citations
- RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4DShuhei Kurita, Naoki Katsura, Eri OnamiICCV 2023 · 26 citations
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 10 citations
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 8 citations
Builds on5
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 251 citations
- A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionYue Liao, Si Liu, Guanbin Li, Fei Wang et al.CVPR 2020
- Memory Enhanced Global-Local Aggregation for Video Object DetectionYihong Chen, Yue Cao, Han Hu, Liwei WangCVPR 2020
- Hypergraph Attention Networks for Multimodal LearningEun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo et al.CVPR 2020
Related papers
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 11 citations
- Language Adaptive Weight Generation for Multi-Task Visual GroundingWei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang et al.CVPR 2023
- Latent Expression Generation for Referring Image Segmentation and GroundingSeonghoon Yu, Joonbeom Hong, Joonseok Lee, Jeany SonICCV 2025 · 1 citation
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.ACM MM 2024 · 6 citations
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu et al.CVPR 2025
