AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe
Abstract
Weakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differences in text expressions due to category-based and attribute-based ambiguity. To address these challenges, we introduce AlignCAT, a novel query-based semantic matching framework for weakly supervised VG. To enhance visual-linguistic alignment, we propose a coarse-grained alignment module that utilizes category information and global context, effectively mitigating interference from category-inconsistent objects. Subsequently, a fine-grained alignment module leverages descriptive information and captures word-level text features to achieve attribute consistency. By exploiting linguistic cues to their fullest extent, our proposed AlignCAT progressively filters out misaligned visual queries and enhances contrastive learning efficiency. Extensive experiments on three VG benchmarks, namely RefCOCO, RefCOCO+, and RefCOCOg, verify the superiority of AlignCAT against existing weakly supervised methods on two VG tasks. Our code is available at: https://github.com/I2-Multimedia-Lab/AlignCAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon et al.CVPR 2022 · 398 citations
- Activation Modulation and Recalibration Scheme for Weakly Supervised Semantic SegmentationJie Qin, Jie Wu, Xuefeng Xiao, Lujun Li et al.AAAI 2022 · 137 citations
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingXuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha et al.ICCV 2019 · 93 citations
- Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language GroundingZhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu et al.NeurIPS 2020 · 74 citations
Related papers
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.ACM MM 2024 · 6 citations
- RefDetector: A Simple Yet Effective Matching-based Method for Referring Expression ComprehensionYabing Wang, Zhuotao Tian, Zheng Qin, Sanping Zhou et al.AAAI 2025 · 2 citations
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 3 citations
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu et al.CVPR 2022 · 146 citations
