RefCLIP: A Universal Teacher for Weakly Supervised Referring Expression Comprehension
Lei Jin, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Annan Shu, Rongrong Ji
摘要
Referring Expression Comprehension (REC) is a task of grounding the referent based on an expression, and its development is greatly limited by expensive instance-level annotations. Most existing weakly supervised methods are built based on two-stage detection networks, which are computationally expensive. In this paper, we resort to the efficient one-stage detector and propose a novel weakly supervised model called RefCLIP. Specifically, RefCLIP redefines weakly supervised REC as an anchor-text matching problem, which can avoid the complex post-processing in existing methods. To achieve weakly supervised learning, we introduce anchor-based contrastive loss to optimize Re-fCLIP via numerous anchor-text pairs. Based on RefCLIP, we further propose the first model-agnostic weakly supervised training scheme for existing REC models, where Ref-CLIP acts as a mature teacher to generate pseudo-labels for teaching common REC models. With our careful designs, this scheme can even help existing REC models achieve better weakly supervised performance than RefCLIP, e.g., TransVG and SimREC. To validate our approaches, we conduct extensive experiments on four REC benchmarks, i.e., RefCOCO, RefCOCO+, RefCOCOg and ReferItGame. Experimental results not only report our significant performance gains over existing weakly supervised models, e.g., +24.87% on RefCOCO, but also show the 5x faster inference speed. Project: https://refclip.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Cycle Consistency as Reward: Learning Image-Text Alignment Without Human PreferencesHyojin Bahng, Caroline Chan, Frédo Durand, Phillip IsolaICCV 2025 · 被引用 25 次
- SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression SegmentationDanni Yang, Jiayi Ji, Yiwei Ma, Tianyu Guo 等ICML 2024 · 被引用 19 次
- Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D RepresentationHaowei Wang, Jiji Tang, Jiayi Ji, Xiaoshuai Sun 等ACM MM 2023 · 被引用 11 次
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 等CVPR 2024 · 被引用 4 次
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unbiased Teacher for Semi-Supervised Object DetectionYen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo 等ICLR 2021 · 被引用 603 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 被引用 317 次
相关 Paper
- ReCLIP: A Strong Zero-Shot Baseline for Referring Expression ComprehensionSanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner 等ACL 2022 · 被引用 172 次
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun 等ACM MM 2024 · 被引用 6 次
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationSilin Cheng, Yang Liu, Xinwei He, Sébastien Ourselin 等CVPR 2025
- RefTeacher: A Strong Baseline for Semi-Supervised Referring Expression ComprehensionJiamu Sun, Gen Luo, Yiyi Zhou, Xiaoshuai Sun 等CVPR 2023
- Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic SegmentationBingfeng Zhang, Siyue Yu, Yunchao Wei, Yao Zhao 等CVPR 2024 · 被引用 35 次
