DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image Captioning
Liangyu Fu, Junbo Wang, Yuke Li, Qiangguo Jin, Hongsong Wang, Jing Ya, Linjiang Huang, Liang Yao, Jiangbin Zheng, Xuecheng Wu, Zhiyong Wang
Abstract
Diffusion-based image captioning methods have been proposed to address the inherent issues of autoregressive models, such as slow inference speed, significant accumulative errors, and limited generative diversity. However, due to excessive reliance on textual data and constrained training objective, existing diffusion-based methods suffer from a semantic gap between vision and language, ultimately resulting in poor quality of generated captions. To address this issue, we propose a novel diffusion-based semantics aligned image captioning framework, namely DSACap. Specifically, DSACap deviates from existing methods which treat text as the target of noise-adding and denoising, instead directly applying these processes to the image, thus reducing the loss of visual-semantic alignment. In addition, we introduce a reinforcement learning-based training strategy to maximize the semantic alignment between image and text. We feed the generated textual descriptions into an image generation model to reconstruct the original image and use the cosine similarity between the generated image and the original image as the reward to train the image captioning model. Extensive experimental results on the MS COCO dataset demonstrate that DSACap achieves a CIDEr score of 128.8, clearly outperforming existing diffusion-based image captioning methods. Our code will be made publicly open soon.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e99ecfc-2cd5-4f8b-96b6-f7e7ea22abf2Related papers
- Semantic-Conditional Diffusion Networks for Image CaptioningJianjie Luo, Yehao Li, Yingwei Pan, Ting Yao et al.CVPR 2023
- Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image CaptioningJiahe Shi, Yali Li, Shengjin WangICCV 2021 · 12 citations
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningLin Zhang, Xianfang Zeng, Kangcong Li, Gang Yu et al.ICCV 2025 · 3 citations
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang et al.ACM MM 2021 · 22 citations
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement LearningHaonan Jia, Shichao Dong, Xin Dong, Zenghui Sun et al.CVPR 2026
