SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
Xiaofu Chen, Israfel Salazar, Yova Kementchedjhieva
Abstract
As interest grows in generating long, detailed image captions, standard evaluation metrics become increasingly unreliable. N-gram-based metrics though efficient, fail to capture semantic correctness. Representational Similarity (RS) metrics, designed to address this, initially saw limited use due to high computational costs, while today, despite advances in hardware, they remain unpopular due to low correlation to human judgments. Meanwhile, metrics based on large language models (LLMs) show strong correlation with human judgments, but remain too expensive for iterative use during model development. We introduce SPECS (Specificity-Enhanced CLIP-Score), a reference-free RS metric tailored to long image captioning. SPECS modifies CLIP with a new objective that emphasizes specificity: rewarding correct details and penalizing incorrect ones. We show that SPECS matches the performance of open-source LLMbased metrics in correlation to human judgments, while being far more efficient. This makes it a practical alternative for iterative checkpoint evaluation during image captioning model development. Our code can be found at https://github.com/mbzuai-nlp/SPECS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a4bcdd2-ef5b-47b8-a7bd-952e51d24d3dBuilds on14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki et al.EMNLP 2025
- Positive-Augmented Contrastive Learning for Image and Video Captioning EvaluationSara Sarto, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi et al.CVPR 2023
- Prompt Expansion for Adaptive Text-to-Image GenerationSiddhartha Datta, Alexander Ku, Deepak Ramachandran, Peter AndersonACL 2024 · 3 citations
