Contrastive Learning with Expectation-Maximization for Weakly Supervised Phrase Grounding
Keqin Chen, Richong Zhang, Samuel Mensah, Yongyi Mao
摘要
Weakly supervised phrase grounding aims to learn an alignment between phrases in a caption and objects in a corresponding image using only caption-image annotations, i.e., without phrase-object annotations. Previous methods typically use a caption-image contrastive loss to indirectly supervise the alignment between phrases and objects, which hinders the maximum use of the intrinsic structure of the multimodal data and leads to unsatisfactory performance. In this work, we directly use the phrase-object contrastive loss in the condition that no positive annotation is available in the first place. Specifically, we propose a novel contrastive learning framework based on the expectation-maximization algorithm that adaptively refines the target prediction. Experiments on two widely used benchmarks, Flickr30K Entities and RefCOCO+, demonstrate the effectiveness of our framework. We obtain 63.05% top-1 accuracy on Flickr30K Entities and 59.51%/43.46% on RefCOCO+ TestA/TestB, outperforming the previous methods by a large margin, even surpassing a previous SoTA that uses a pre-trained vision-language model. Furthermore, we deliver a theoretical analysis of the effectiveness of our method from the perspective of the maximum likelihood estimate with latent variables.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning to Segment Referred Objects from Narrated Egocentric VideosYuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli 等CVPR 2024 · 被引用 1 次
- Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase LocalizationPengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia 等AAAI 2026
- Momentum Pseudo-Labeling for Weakly Supervised Phrase GroundingDongdong Kuang, Richong Zhang, Zhijie Nie, Junfan Chen 等AAAI 2025
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationLiwei Wang, Jing Huang, Yin Li, Kun Xu 等CVPR 2021
- AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingYidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 等ACM MM 2025 · 被引用 2 次
- Improving Visual Grounding by Encouraging Consistent Gradient-Based ExplanationsZiyan Yang, Kushal Kafle, Franck Dernoncourt, Vicente OrdonezCVPR 2023
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao 等ACM MM 2020 · 被引用 14 次
