Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation
Sara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
Abstract
The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positive-Augmented Contrastive learning Score (PAC-S), that in a novel way unifies the learning of a contrastive visual-semantic space with the addition of generated images and text on curated data. Experiments spanning several datasets demonstrate that our new metric achieves the highest correlation with human judgments on both images and videos, outperforming existing referencebased metrics like CIDEr and SPICE and reference-free metrics like CLIP-Score. Finally, we test the system-level correlation of the proposed metric when considering popular image captioning approaches, and assess the impact of employing different cross-modal features. Our source code and trained models are publicly available at: https: //github.com/aimagelab/pacscore .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a412202a-12ee-4e3d-8827-d1e8cadc0404Cited by top-tier papers26
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4oTony Cheng Tong, Sirui He, Zhiwen Shao, Dit-Yan YeungAAAI 2025 · 22 citations
- Polos: Multimodal Metric Learning from Human Feedback for Image CaptioningYuiga Wada, Kanta Kaneda, Daichi Saito, Komei SugiuraCVPR 2024 · 16 citations
- ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language ModelsDuy M. H. Nguyen, Nghiem Tuong Diep, Trung Nguyen, Hoang-Bao Le et al.NeurIPS 2025 · 7 citations
- Panoptic Captioning: An Equivalence Bridge for Image and TextKun-Yu Lin, Hongjun Wang, Weining Ren, Kai HanNeurIPS 2025 · 7 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- Mutual Information Divergence: A Unified Metric for Multimodal Generative ModelsJin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo et al.NeurIPS 2022 · 49 citations
- RCA-NOC: Relative Contrastive Alignment for Novel Object CaptioningJiashuo Fan, Yaoyuan Liang, Leyao Liu, Shao-Lun Huang et al.ICCV 2023 · 7 citations
- CONICA: A Contrastive Image Captioning Framework with Robust Similarity LearningLin Deng, Yuzhong Zhong, Maoning Wang, Jianwei ZhangACM MM 2023 · 4 citations
