ReFormer: The Relational Transformer for Image Captioning
Xuewen Yang, Yingru Liu, Xin Wang
摘要
Image captioning is shown to be able to achieve a better performance by using scene graphs to represent the relations of objects in the image. The current captioning encoders generally use a Graph Convolutional Net (GCN) to represent the relation information and merge it with the object region features via concatenation or convolution to get the final input for sentence decoding. However, the GCN-based encoders in the existing methods are less effective for captioning due to two reasons. First, using the image captioning as the objective (i.e., Maximum Likelihood Estimation) rather than a relation-centric loss cannot fully explore the potential of the encoder. Second, using a pre-trained model instead of the encoder itself to extract the relationships is not flexible and cannot contribute to the explainability of the model. To improve the quality of image captioning, we propose a novel architecture ReFormer- a RElational transFORMER to generate features with relation information embedded and to explicitly express the pair-wise relationships between objects in the image. ReFormer incorporates the objective of scene graph generation with that of image captioning using one modified Transformer model. This design allows ReFormer to generate not only better image captions with the benefit of extracting strong relational image features, but also scene graphs to explicitly describe the pair-wise relationships. Experiments on publicly available datasets show that our model significantly outperforms state-of-the-art methods on image captioning and scene graph generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 被引用 108 次
- Visually-Prompted Language Model for Fine-Grained Scene Graph Generation in an Open WorldQifan Yu, Juncheng Li, Yu Wu, Siliang Tang 等ICCV 2023 · 被引用 51 次
- Efficient Image Captioning for Edge DevicesNing Wang, Jiangrong Xie, Hang Luo, Qinglin Cheng 等AAAI 2023 · 被引用 41 次
- Uncertainty-Aware Image CaptioningZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 等AAAI 2023 · 被引用 21 次
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song 等EMNLP 2023 · 被引用 16 次
它引用的顶会 Paper8
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Learning Tuple Compatibility for Conditional Outfit RecommendationXuewen Yang, Dongliang Xie, Xin Wang, Jiangbo Yuan 等ACM MM 2020 · 被引用 25 次
- Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task LearningYingru Liu, Xuewen Yang, Dongliang Xie, Xin Wang 等AAAI 2020 · 被引用 10 次
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
相关 Paper
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningXinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia XiaoACM MM 2021 · 被引用 75 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha 等ICCV 2021 · 被引用 55 次
- Storytelling from an Image Stream Using Scene GraphsRuize Wang, Zhongyu Wei, Piji Li, Qi Zhang 等AAAI 2020 · 被引用 75 次
- Transforming Visual Scene Graphs to Image CaptionsXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu 等ACL 2023 · 被引用 22 次
