In Defense of Scene Graphs for Image Captioning
Kien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha, Truong Q. Nguyen
Abstract
The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics, such as object entities, relationships and attributes. Several studies have noted that the naive use of scene graphs from a black-box scene graph generator harms image captioning performance and that scene graph-based captioning models have to incur the overhead of explicit use of image features to generate decent captions. Addressing these challenges, we propose SG2Caps, a framework that utilizes only the scene graph labels for competitive image captioning performance. The basic idea is to close the semantic gap between the two scene graphs - one derived from the input image and the other from its caption. In order to achieve this, we leverage the spatial location of objects and the Human-Object-Interaction (HOI) labels as an additional HOI graph. SG2Caps outperforms existing scene graph-only captioning models by a large margin, indicating scene graphs as a promising representation for image captioning. Direct utilization of scene graph labels avoids expensive graph convolutions over high-dimensional CNN features resulting in 49% fewer trainable parameters. Our code is available at: https://github.com/Kien085/SG2Caps
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Composing Object Relations and Attributes for Image-Text MatchingKhoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2024 · 31 citations
- Transforming Visual Scene Graphs to Image CaptionsXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu et al.ACL 2023 · 22 citations
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi et al.CVPR 2024 · 14 citations
- Multiview Scene GraphJuexiao Zhang, Gao Zhu, Sihang Li, Xinhao Liu et al.NeurIPS 2024 · 13 citations
Builds on4
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph GenerationShaotian Yan, Chen Shen, Zhongming Jin, Jianqiang Huang et al.ACM MM 2020 · 115 citations
- VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph ConvolutionsOytun Ulutan, A. S. M. Iftekhar, B. S. ManjunathCVPR 2020
- Unbiased Scene Graph Generation From Biased TrainingKaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi et al.CVPR 2020
Related papers
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 70 citations
- Hierarchical Scene Graph Encoder-Decoder for Image Paragraph CaptioningXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiACM MM 2020 · 25 citations
- Say As You Wish: Fine-Grained Control of Image Caption Generation With Abstract Scene GraphsShizhe Chen, Qin Jin, Peng Wang, Qi WuCVPR 2020
- Learning to Generate Scene Graph from Natural Language SupervisionYiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu et al.ICCV 2021 · 88 citations
- Linguistic Structures As Weak Supervision for Visual Scene Graph GenerationKeren Ye, Adriana KovashkaCVPR 2021
