Similarity Maps for Self-Training Weakly-Supervised Phrase Grounding
Tal Shaharabany, Lior Wolf
摘要
A phrase grounding model receives an input image and a text phrase and outputs a suitable localization map. We present an effective way to refine a phrase ground model by considering self-similarity maps extracted from the latent representation of the model's image encoder. Our main insights are that these maps resemble localization maps and that by combining such maps, one can obtain useful pseudo-labels for performing self-training. Our results surpass, by a large margin, the state of the art in weakly supervised phrase grounding. A similar gap in performance is obtained for a recently proposed downstream task called WWbL, in which only the image is input, without any text. Our code is available at https://github .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 等CVPR 2024 · 被引用 4 次
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingTa Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi 等ICCV 2025 · 被引用 2 次
- Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase LocalizationPengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia 等AAAI 2026
- Momentum Pseudo-Labeling for Weakly Supervised Phrase GroundingDongdong Kuang, Richong Zhang, Zhijie Nie, Junfan Chen 等AAAI 2025
- Improved Visual Grounding through Self-Consistent ExplanationsRuozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C. Berg 等CVPR 2024
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano 等SIGGRAPH 2022 · 被引用 501 次
相关 Paper
- What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsTal Shaharabany, Yoad Tewel, Lior WolfNeurIPS 2022 · 被引用 26 次
- Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption AlignmentSamyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja 等ICCV 2019 · 被引用 113 次
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationLiwei Wang, Jing Huang, Yin Li, Kun Xu 等CVPR 2021
- Confidence-aware Pseudo-label Learning for Weakly Supervised Visual GroundingYang Liu, Jiahua Zhang, Qingchao Chen, Yuxin PengICCV 2023 · 被引用 19 次
- Box-based Refinement for Weakly Supervised and Unsupervised Localization TasksEyal Gomel, Tal Shaharabany, Lior WolfICCV 2023 · 被引用 6 次
