Correspondence Matters for Video Referring Expression Comprehension
Meng Cao, Ji Jiang, Long Chen, Yuexian Zou
Abstract
We investigate the problem of video Referring Expression Comprehension (REC), which aims to localize the referent objects described in the sentence to visual regions in the video frames. Despite the recent progress, existing methods suffer from two problems: 1) inconsistent localization results across video frames; 2) confusion between the referent and contextual objects. To this end, we propose a novel Dual Correspondence Network (dubbed as DCNet) which explicitly enhances the dense associations in both the inter-frame and cross-modal manners. Firstly, we aim to build the inter-frame correlations for all existing instances within the frames. Specifically, we compute the inter-frame patch-wise cosine similarity to estimate the dense alignment and then perform the inter-frame contrastive learning to map them close in feature space. Secondly, we propose to build the fine-grained patch-word alignment to associate each patch with certain words. Due to the lack of this kind of detailed annotations, we also predict the patch-word correspondence through the cosine similarity. Extensive experiments demonstrate that our DCNet achieves state-of-the-art performance on both video and image REC benchmarks. Furthermore, we conduct comprehensive ablation studies and thorough analyses to explore the optimal model designs. Notably, our inter-frame and cross-modal contrastive losses are plug-and-play functions and are applicable to any video REC architectures. For example, by building on top of Co-grounding [51], we boost the performance by 1.48% absolute improvement on Accu.@0.5 for VID-Sentence dataset. Our codes are available at https://github.com/mengcaopku/DCNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eca0c2ce-3b96-4b0f-bf1f-d190f4aa7974Cited by top-tier papers6
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.ICCV 2023 · 32 citations
- Open-Vocabulary Object Detection via Scene Graph DiscoveryHengcan Shi, Munawar Hayat, Jianfei CaiACM MM 2023 · 20 citations
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video GroundingYang Zhan, Yuan YuanNeurIPS 2025 · 8 citations
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video UnderstandingPengfei Hu, Meng Cao, Yingyao Wang, Yi Wang et al.CVPR 2026 · 3 citations
- Iterative Proposal Refinement for Weakly-Supervised Video GroundingMeng Cao, Fangyun Wei, Can Xu, Xiubo Geng et al.CVPR 2023
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
Related papers
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan et al.SIGIR 2021 · 88 citations
- Multi-Task Collaborative Network for Joint Referring Expression Comprehension and SegmentationGen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao et al.CVPR 2020
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 6 citations
- Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in VideosSijie Song, Xudong Lin, Jiaying Liu, Zongming Guo et al.CVPR 2021
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 21 citations
