Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image Captioning
Weizhi Nie, Jiesi Li, Ning Xu, An-An Liu, Xuanya Li, Yongdong Zhang
摘要
Image captioning aims to generate a sentence consisting of sequential linguistic words, to describe visual units (i.e., objects, relationships, and attributes) in a given image. Most of existing methods rely on the prevalent supervised learning with cross-entropy (XE) function to transfer visual units into a sequence of linguistic words. However, we argue that the XE objective is not sensitive to visual-linguistic alignment, which cannot discriminately penalize the semantic inconsistency and shrink the context gap. To solve these problems, we propose the Triangle-Reward Reinforcement Learning (TRRL) method. TRRL uses the scene graph (G)---objects as nodes and relationships as edges---to represent images, generated sentences, and ground truth sentences individually, and mutually align them during the training process. Specifically, TRRL formulates the image captioning into cooperative agents, where the first agent aims to extract visual scene graph (Gimg) from image (I) and the second agent translates this graph into sentence (S). To discriminately penalize the visual-linguistic inconsistency, TRRL proposes the novel triangle-reward function: 1) the generated sentence and its corresponding ground truth are decomposed into the linguistic scene graph (Gsen) and ground-truth scene graph (Ggt), respectively; 2) Gimg, Gsen, and Ggt are paired to calculate the semantic similarity scores which are proportionally assigned to reward each agent. Meanwhile, to make the training objective sensitive to context changes, we propose the node-level and triplet-level scoring methods to jointly measure the visual-linguistic graph correlations. Extensive experiments on the MSCOCO dataset demonstrate the superiority of TRRL. Additional ablation studies further validate its effectiveness.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Progressive Tree-Structured Prototype Network for End-to-End Image CaptioningPengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli GaoACM MM 2022 · 被引用 35 次
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song 等EMNLP 2023 · 被引用 16 次
- VinaBench: Benchmark for Faithful and Consistent Visual NarrativesSilin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler 等CVPR 2025
相关 Paper
- Multi-modal Dependency Tree for Video CaptioningWentian Zhao, Xinxiao Wu, Jiebo LuoNeurIPS 2021 · 被引用 27 次
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningLin Zhang, Xianfang Zeng, Kangcong Li, Gang Yu 等ICCV 2025 · 被引用 3 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
- DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image CaptioningLiangyu Fu, Junbo Wang, Yuke Li, Qiangguo Jin 等ACM MM 2025
