CONICA: A Contrastive Image Captioning Framework with Robust Similarity Learning
Lin Deng, Yuzhong Zhong, Maoning Wang, Jianwei Zhang
Abstract
Contrastive Language Image Pre-training (CLIP) has recently made significant advancements in image captioning by providing effective multi-modal representation learning capabilities. However, previous studies primarily rely on the language-aligned visual semantics as input for the captioning model, leaving the learned robust vision-language relevance under-exploited. In this paper, we propose CONICA, a unified CONtrastive Image CAptioning framework that investigates how contrastive learning can further enhance image captioning from three aspects. Firstly, we introduce contrastive learning objectives into the typical image captioning training pipeline with minimal overhead. Secondly, we construct fine-grained contrastive samples to obtain image-text similarities that correlate with the evaluation metric of image captioning. Finally, we incorporate the learned contrastive knowledge into the captioning decoding strategy to search for better captions. Experimental results demonstrate that CONICA significantly improves performance over standard captioning baselines and achieves new state-of-the-art results on the MSCOCO and Flikr30K. Source code is available at https://github.com/DenglinGo/CONICA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal AlignmentZiping Ma, Furong Xu, Jian Liu, Ming Yang et al.ICML 2024 · 8 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- RCA-NOC: Relative Contrastive Alignment for Novel Object CaptioningJiashuo Fan, Yaoyuan Liang, Leyao Liu, Shao-Lun Huang et al.ICCV 2023 · 7 citations
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
- CgT-GAN: CLIP-guided Text GAN for Image CaptioningJiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu et al.ACM MM 2023 · 26 citations
