Generating Diverse and Descriptive Image Captions Using Visual Paraphrases
Lixin Liu, Jiajun Tang, Xiaojun Wan, Zongming Guo
Abstract
Recently there has been significant progress in image captioning with the help of deep learning. However, captions generated by current state-of-the-art models are still far from satisfactory, despite high scores in terms of conventional metrics such as BLEU and CIDEr. Human-written captions are diverse, informative and precise, but machine-generated captions seem to be simple, vague and dull. In this paper, aimed at improving diversity and descriptiveness characteristics of generated image captions, we propose a model utilizing visual paraphrases (different sentences describing the same image) in captioning datasets. We explore different strategies to select useful visual paraphrase pairs for training by designing a variety of scoring functions. Our model consists of two decoding stages, where a preliminary caption is generated in the first stage and then paraphrased into a more diverse and descriptive caption in the second stage. Extensive experiments are conducted on the benchmark MS COCO dataset, with automatic evaluation and human evaluation results verifying the effectiveness of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0d48366-ec3c-4cf9-8096-e590d2d114a3Cited by top-tier papers7
- Rethinking the Reference-based Distinctive Image CaptioningYangjun Mao, Long Chen, Zhihong Jiang, Dong Zhang et al.ACM MM 2022 · 22 citations
- CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained KnowledgeLinli Yao, Weijing Chen, Qin JinWWW 2023 · 11 citations
- Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood EstimationZihao Yue, Anwen Hu, Liang Zhang, Qin JinNeurIPS 2023 · 7 citations
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationJun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon et al.EMNLP 2023
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki et al.EMNLP 2025
Related papers
- Bridging the Gap between Vision and Language Domains for Improved Image CaptioningFenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang et al.ACM MM 2020 · 13 citations
- Paraphrasing Is All You Need for Novel Object CaptioningCheng-Fu Yang, Yao-Hung Hubert Tsai, Wan-Cyuan Fan, Russ Salakhutdinov et al.NeurIPS 2022 · 7 citations
- Image Captioning with Multimodal Guidance and Search Space OptimizationYimou Guo, Yaochen Li, Jingze Liu, Jiahui Feng et al.ACM MM 2025
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 36 citations
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen et al.ICCV 2019 · 107 citations
