Reinforcing an Image Caption Generator Using Off-Line Human Feedback
Paul Hongsuck Seo, Piyush Sharma, Tomer Levinboim, Bohyung Han, Radu Soricut
Abstract
Human ratings are currently the most accurate way to assess the quality of an image captioning model, yet most often the only used outcome of an expensive human rating evaluation is a few overall statistics over the evaluation dataset. In this paper, we show that the signal from instance-level human caption ratings can be leveraged to improve captioning models, even when the amount of caption ratings is several orders of magnitude less than the caption training data. We employ a policy gradient method to maximize the human ratings as rewards in an off-policy reinforcement learning setting, where policy gradients are estimated by samples from a distribution that focuses on the captions in a caption ratings dataset. Our empirical evidence indicates that the proposed method learns to generalize the human raters' judgments to a previously unseen set of images, as judged by a different set of human judges, and additionally on a different, multi-dimensional side-by-side human evaluation procedure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 653d51d2-c001-4181-896f-f9b99eb9ee2eCited by top-tier papers3
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- Look Before You Speak: Visually Contextualized UtterancesPaul Hongsuck Seo, Arsha Nagrani, Cordelia SchmidCVPR 2021
- Affection: Learning Affective Explanations for Real-World Visual DataPanos Achlioptas, Maks Ovsjanikov, Leonidas J. Guibas, Sergey TulyakovCVPR 2023
Related papers
- Partial Off-policy Learning: Balance Accuracy and Diversity for Human-Oriented Image CaptioningJiahe Shi, Yali Li, Shengjin WangICCV 2021 · 12 citations
- CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image CaptioningZhijiang Tang, Linhua Wang, Jiaxin Qi, Weihao Jiang et al.CVPR 2026 · 7 citations
- CapWAP: Image Captioning with a PurposeAdam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark et al.EMNLP 2020 · 17 citations
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao et al.ICLR 2026 · 37 citations
