Relation-aware Instance Refinement for Weakly Supervised Visual Grounding
Yongfei Liu, Bo Wan, Lin Ma, Xuming He
Abstract
Visual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize weak supervision from only image-caption pairs. Previous methods typically rely on matching query phrases directly to a precomputed, fixed object candidate pool, which leads to inaccurate localization and ambiguous matching due to lack of semantic relation constraints.
In our paper, we propose a novel context-aware weaklysupervised learning method that incorporates coarse-tofine object refinement and entity relation modeling into a two-stage deep network, capable of producing more accurate object representation and matching. To effectively train our network, we introduce a self-taught regression loss for the proposal locations and a classification loss based on parsed entity relations. Extensive experiments on two public benchmarks Flickr30K Entities and Refer-ItGame demonstrate the efficacy of our weakly grounding framework. The results show that we outperform the previous methods by a considerable margin, achieving 59.27% top-1 accuracy in Flickr30K Entities and 37.68% in the ReferItGame dataset respectively 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76b065fd-c1d8-40ae-8548-cf6da5564c92Cited by top-tier papers21
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon et al.CVPR 2022 · 398 citations
- TubeDETR: Spatio-Temporal Video Grounding with TransformersAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.CVPR 2022 · 87 citations
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu et al.ICCV 2023 · 52 citations
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 29 citations
Builds on9
- Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion NetworkBairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang et al.ICCV 2019 · 183 citations
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja et al.ICCV 2019 · 113 citations
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingXuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha et al.ICCV 2019 · 93 citations
- Phrase Localization Without Paired Training ExamplesJosiah Wang, Lucia SpeciaICCV 2019 · 51 citations
Related papers
- Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase GroundingKeqin Chen, Richong Zhang, Samuel Mensah, Yongyi MaoEMNLP 2022 · 3 citations
- Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.ICCV 2023 · 30 citations
- Confidence-aware Pseudo-label Learning for Weakly Supervised Visual GroundingYang Liu, Jiahua Zhang, Qingchao Chen, Yuxin PengICCV 2023 · 19 citations
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.ACM MM 2024 · 6 citations
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
