Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in Videos
Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo, Shih-Fu Chang
摘要
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking, proposal-based matching), we tackle the problem from a novel perspective, co-grounding, with an elegant one-stage framework. We enhance the single-frame grounding accuracy by semantic attention learning and improve the cross-frame grounding consistency with co-grounding feature learning. Semantic attention learning explicitly parses referring cues in different attributes to reduce the ambiguity in the complex expression. Co-grounding feature learning boosts visual feature representations by integrating temporal correlation to reduce the ambiguity caused by scene dynamics. Experiment results demonstrate the superiority of our framework on the video grounding datasets VID and LiOTB in generating accurate and stable results across frames. Our model is also applicable to referring expression comprehension in images, illustrated by the improved performance on the RefCOCO dataset. Our project is available at https://sijiesong.github.io/co- grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- Described Object Detection: Liberating Object Detection with Flexible ExpressionsChi Xie, Zhao Zhang, Yixuan Wu, Feng Zhu 等NeurIPS 2023 · 被引用 69 次
- RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4DShuhei Kurita, Naoki Katsura, Eri OnamiICCV 2023 · 被引用 26 次
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 被引用 10 次
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper5
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 被引用 251 次
- A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionYue Liao, Si Liu, Guanbin Li, Fei Wang 等CVPR 2020
- Memory Enhanced Global-Local Aggregation for Video Object DetectionYihong Chen, Yue Cao, Han Hu, Liwei WangCVPR 2020
- Hypergraph Attention Networks for Multimodal LearningEun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo 等CVPR 2020
相关 Paper
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 被引用 11 次
- Language Adaptive Weight Generation for Multi-Task Visual GroundingWei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang 等CVPR 2023
- Latent Expression Generation for Referring Image Segmentation and GroundingSeonghoon Yu, Joonbeom Hong, Joonseok Lee, Jeany SonICCV 2025 · 被引用 1 次
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun 等ACM MM 2024 · 被引用 6 次
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 等CVPR 2025
