Correspondence Matters for Video Referring Expression Comprehension
Meng Cao, Ji Jiang, Long Chen, Yuexian Zou
摘要
We investigate the problem of video Referring Expression Comprehension (REC), which aims to localize the referent objects described in the sentence to visual regions in the video frames. Despite the recent progress, existing methods suffer from two problems: 1) inconsistent localization results across video frames; 2) confusion between the referent and contextual objects. To this end, we propose a novel Dual Correspondence Network (dubbed as DCNet) which explicitly enhances the dense associations in both the inter-frame and cross-modal manners. Firstly, we aim to build the inter-frame correlations for all existing instances within the frames. Specifically, we compute the inter-frame patch-wise cosine similarity to estimate the dense alignment and then perform the inter-frame contrastive learning to map them close in feature space. Secondly, we propose to build the fine-grained patch-word alignment to associate each patch with certain words. Due to the lack of this kind of detailed annotations, we also predict the patch-word correspondence through the cosine similarity. Extensive experiments demonstrate that our DCNet achieves state-of-the-art performance on both video and image REC benchmarks. Furthermore, we conduct comprehensive ablation studies and thorough analyses to explore the optimal model designs. Notably, our inter-frame and cross-modal contrastive losses are plug-and-play functions and are applicable to any video REC architectures. For example, by building on top of Co-grounding [51], we boost the performance by 1.48% absolute improvement on Accu.@0.5 for VID-Sentence dataset. Our codes are available at https://github.com/mengcaopku/DCNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li 等ICCV 2023 · 被引用 32 次
- Open-Vocabulary Object Detection via Scene Graph DiscoveryHengcan Shi, Munawar Hayat, Jianfei CaiACM MM 2023 · 被引用 20 次
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 被引用 8 次
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video UnderstandingPengfei Hu, Meng Cao, Yingyao Wang, Yi Wang 等CVPR 2026 · 被引用 3 次
- Iterative Proposal Refinement for Weakly-Supervised Video GroundingMeng Cao, Fangyun Wei, Can Xu, Xiubo Geng 等CVPR 2023
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 被引用 317 次
相关 Paper
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan 等SIGIR 2021 · 被引用 88 次
- Multi-Task Collaborative Network for Joint Referring Expression Comprehension and SegmentationGen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao 等CVPR 2020
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 被引用 6 次
- Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in VideosSijie Song, Xudong Lin, Jiaying Liu, Zongming Guo 等CVPR 2021
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 被引用 21 次
