Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image Captioning
Jiahe Shi, Yali Li, Shengjin Wang
Abstract
Human-oriented image captioning with both high diversity and accuracy is a challenging task in vision+language modeling. The reinforcement learning (RL) based frameworks promote the accuracy of image captioning, yet seriously hurt the diversity. In contrast, other methods based on variational auto-encoder (VAE) or generative adversarial network (GAN) can produce diverse yet less accurate captions. In this work, we devote our attention to promote the diversity of RL-based image captioning. To be specific, we devise a partial off-policy learning scheme to balance accuracy and diversity. First, we keep the model exposed to varied candidate captions by sampling from the initial state before RL launched. Second, a novel criterion named max-CIDEr is proposed to serve as the reward for promoting diversity. We combine the above-mentioned offpolicy strategy with the on-policy one to moderate the exploration effect, further balancing the diversity and accuracy for human-like image captioning. Experiments show that our method locates the closest to human performance in the diversity-accuracy space, and achieves the highest Pearson correlation as 0.337 with human performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2d939d9-9e27-491c-84bb-6d1b59bebf34Builds on4
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 71 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- Better Captioning With Sequence-Level ExplorationJia Chen, Qin JinCVPR 2020
Related papers
- CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image CaptioningZhijiang Tang, Linhua Wang, Jiaxin Qi, Weihao Jiang et al.CVPR 2026 · 7 citations
- Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text GenerationWenqing Chen, Jidong Tian, Caoyun Fan, Yitian Li et al.AAAI 2023 · 2 citations
- DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image CaptioningLiangyu Fu, Junbo Wang, Yuke Li, Qiangguo Jin et al.ACM MM 2025
- Reinforcing an Image Caption Generator Using Off-Line Human FeedbackPaul Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han et al.AAAI 2020 · 24 citations
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 48 citations
