Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image Captioning
Weizhi Nie, Jiesi Li, Ning Xu, An-An Liu, Xuanya Li, Yongdong Zhang
Abstract
Image captioning aims to generate a sentence consisting of sequential linguistic words, to describe visual units (i.e., objects, relationships, and attributes) in a given image. Most of existing methods rely on the prevalent supervised learning with cross-entropy (XE) function to transfer visual units into a sequence of linguistic words. However, we argue that the XE objective is not sensitive to visual-linguistic alignment, which cannot discriminately penalize the semantic inconsistency and shrink the context gap. To solve these problems, we propose the Triangle-Reward Reinforcement Learning (TRRL) method. TRRL uses the scene graph (G)---objects as nodes and relationships as edges---to represent images, generated sentences, and ground truth sentences individually, and mutually align them during the training process. Specifically, TRRL formulates the image captioning into cooperative agents, where the first agent aims to extract visual scene graph (Gimg) from image (I) and the second agent translates this graph into sentence (S). To discriminately penalize the visual-linguistic inconsistency, TRRL proposes the novel triangle-reward function: 1) the generated sentence and its corresponding ground truth are decomposed into the linguistic scene graph (Gsen) and ground-truth scene graph (Ggt), respectively; 2) Gimg, Gsen, and Ggt are paired to calculate the semantic similarity scores which are proportionally assigned to reward each agent. Meanwhile, to make the training objective sensitive to context changes, we propose the node-level and triplet-level scoring methods to jointly measure the visual-linguistic graph correlations. Extensive experiments on the MSCOCO dataset demonstrate the superiority of TRRL. Additional ablation studies further validate its effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 821a14dd-5e1e-47d2-b8f8-90fe40ef6c61Cited by top-tier papers3
- Progressive Tree-Structured Prototype Network for End-to-End Image CaptioningPengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli GaoACM MM 2022 · 35 citations
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song et al.EMNLP 2023 · 16 citations
- VinaBench: Benchmark for Faithful and Consistent Visual NarrativesSilin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler et al.CVPR 2025
Related papers
- Multi-modal Dependency Tree for Video CaptioningWentian Zhao, Xinxiao Wu, Jiebo LuoNeurIPS 2021 · 27 citations
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningLin Zhang, Xianfang Zeng, Kangcong Li, Gang Yu et al.ICCV 2025 · 3 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li et al.CVPR 2020
- DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image CaptioningLiangyu Fu, Junbo Wang, Yuke Li, Qiangguo Jin et al.ACM MM 2025
