PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
Ming Dai, Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng, Wankou Yang
Abstract
Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the model's capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses objectand semantic-level information to improve the recognition of absent targets. Extensive experiments on gRe-fCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO/+/g, and RefCOCO/+/g (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cac90bde-7462-47de-b2b0-5eb02d930f42Cited by top-tier papers3
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback SynergyMing Dai, Wenxuan Cheng, Jiang-Jiang Liu, Sen Yang et al.ICCV 2025 · 5 citations
- VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object SegmentationMing Dai, Sen Yang, Boqiang Duan, Boyuan Tong et al.ICML 2026
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
Builds on42
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
Related papers
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
- Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsMing Dai, Jian Li, Jiedong Zhuang, Xian Zhang et al.AAAI 2025 · 23 citations
- Latent Expression Generation for Referring Image Segmentation and GroundingSeonghoon Yu, Joonbeom Hong, Joonseok Lee, Jeany SonICCV 2025 · 1 citation
- EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual GroundingGwangWook Park, Hyo-Jun Lee, Jong-Hyeon Baek, Hanul Kim et al.CVPR 2026
- CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual GroundingJianjun Zhang, Hanli WangACL 2026
