ReFormer: The Relational Transformer for Image Captioning
Xuewen Yang, Yingru Liu, Xin Wang
Abstract
Image captioning is shown to be able to achieve a better performance by using scene graphs to represent the relations of objects in the image. The current captioning encoders generally use a Graph Convolutional Net (GCN) to represent the relation information and merge it with the object region features via concatenation or convolution to get the final input for sentence decoding. However, the GCN-based encoders in the existing methods are less effective for captioning due to two reasons. First, using the image captioning as the objective (i.e., Maximum Likelihood Estimation) rather than a relation-centric loss cannot fully explore the potential of the encoder. Second, using a pre-trained model instead of the encoder itself to extract the relationships is not flexible and cannot contribute to the explainability of the model. To improve the quality of image captioning, we propose a novel architecture ReFormer- a RElational transFORMER to generate features with relation information embedded and to explicitly express the pair-wise relationships between objects in the image. ReFormer incorporates the objective of scene graph generation with that of image captioning using one modified Transformer model. This design allows ReFormer to generate not only better image captions with the benefit of extracting strong relational image features, but also scene graphs to explicitly describe the pair-wise relationships. Experiments on publicly available datasets show that our model significantly outperforms state-of-the-art methods on image captioning and scene graph generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang et al.ICCV 2023 · 51 citations
- Efficient Image Captioning for Edge DevicesNing Wang, Jiangrong Xie, Hang Luo, Qinglin Cheng et al.AAAI 2023 · 41 citations
- Uncertainty-Aware Image CaptioningZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang et al.AAAI 2023 · 21 citations
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song et al.EMNLP 2023 · 16 citations
Builds on8
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Learning Tuple Compatibility for Conditional Outfit RecommendationXuewen Yang, Dongliang Xie, Xin Wang, Jiangbo Yuan et al.ACM MM 2020 · 25 citations
- Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task LearningYingru Liu, Xuewen Yang, Dongliang Xie, Xin Wang et al.AAAI 2020 · 10 citations
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
Related papers
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningXinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia XiaoACM MM 2021 · 75 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang et al.AAAI 2020 · 75 citations
- Transforming Visual Scene Graphs to Image CaptionsXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu et al.ACL 2023 · 22 citations
