QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding
Shengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Rongrong Ji
Abstract
Visual grounding is a task of locating the object referred by a natural language description. To reduce annotation costs, recent researchers are devoted into one-stage weakly supervised methods for visual grounding, which typically adopt the anchor-text matching paradigm. Despite the efficiency, we identify that anchor representations are often noisy and insufficient to describe object information, which inevitably hinders the vision-language alignments. In this paper, we propose a novel query-based one-stage framework for weakly supervised visual grounding, namely QueryMatch. Different from previous work, QueryMatch represents candidate objects with a set of query features, which inherently establish accurate one-to-one associations with visual objects. In this case, QueryMatch re-formulates weakly supervised visual grounding as a query-text matching problem, which can be optimized via the query-based contrastive learning. Based on QueryMatch, we further propose an innovative strategy for effective weakly supervised learning, namely Active Query Selection (AQS). In particular, AQS aims to enhance the effectiveness of query-based contrastive learning by actively selecting high-quality query features. Through this strategy, AQS can greatly benefit the weakly supervised learning of QueryMatch. To validate our approach, we conduct extensive experiments on three benchmark datasets of two grounding tasks, i.e., referring expression comprehension (REC) and segmentation (RES). Experimental results not only show the state-of-art performance of QueryMatch in two tasks, e.g., over +5% [email protected] on RefCOCO in REC and over +20% mIOU on RefCOCO in RES, but also confirm the effectiveness of AQS in weakly supervised learning. Source codes are available at https://github.com/TensorThinker/QueryMatch.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 316e518d-a3c4-4d71-99ea-52422033aaebCited by top-tier papers3
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao et al.ACM MM 2025 · 2 citations
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingTa Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi et al.ICCV 2025 · 2 citations
- Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual GroundingYuzhen Li, Min Liu, Yuan Bian, Xueping Wang et al.ACM MM 2025 · 1 citation
Related papers
- One-Stage Visual Grounding via Semantic-Aware Feature FilterJiabo Ye, Xin Lin, Liang He, Dingbang Li et al.ACM MM 2021 · 38 citations
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- RefCLIP: A Universal Teacher for Weakly Supervised Referring Expression ComprehensionLei Jin, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.CVPR 2023
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationSilin Cheng, Yang Liu, Xinwei He, Sébastien Ourselin et al.CVPR 2025
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
