SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
Lin Zhang, Xianfang Zeng, Kangcong Li, Gang Yu, Tao Chen
摘要
We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decomposed into object, attribute, and relation sets using scene-graph parsing algorithms. We calculate the set difference between sets of initial and self-corrected captions to identify added and removed elements. These elements are matched against the reference sets to calculate correctness bonuses for accurate refinements and mistake punishments for wrong additions and removals, thereby forming the final reward. For image caption quality assessment, we propose a set of metrics refined from CAPTURE that alleviate its incomplete precision evaluation and inefficient relation matching problems. Furthermore, we collect a fine-grained annotated image caption dataset, RefinedCaps, consisting of 6.5K diverse images from COCO dataset. Experiments show that applying SC-Captioner on large visual-language models can generate better image captions across various scenarios, significantly outperforming the direct preference optimization training strategy. Our code
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG 等CVPR 2026 · 被引用 7 次
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source SuitesZhenxin Lei, Zhangwei Gao, Changyao Tian, Erfei Cui 等ICLR 2026 · 被引用 1 次
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement LearningHaonan Jia, Shichao Dong, Xin Dong, Zenghui Sun 等CVPR 2026
- HSGG: Training-Free Hierarchical Scene Graph Generation with Geometry-Guided Relation Reasoningyunzhe Liu, Wenbiao Liu, Lihui Cen, Zhe Qu 等ICML 2026
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake DetectionTuan Nguyen, Naseem Khan, Khang Tran, Hai Phan 等ICML 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image CaptioningZhijiang Tang, Linhua Wang, Jiaxin Qi, Weihao Jiang 等CVPR 2026 · 被引用 7 次
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu 等ACM MM 2021 · 被引用 9 次
- DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image CaptioningLiangyu Fu, Junbo Wang, Yuke Li, Qiangguo Jin 等ACM MM 2025
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 被引用 70 次
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 等ICLR 2026 · 被引用 37 次
