Transforming Visual Scene Graphs to Image Captions
Xu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, Songfang Huang, Fei Huang, Zhangzikang Li, Yu Zhang
Abstract
We propose to TransForm Scene Graphs into more descriptive Captions (TFSGC). In TF-SGC, we apply multi-head attention (MHA) to design the Graph Neural Network (GNN) for embedding scene graphs. After embedding, different graph embeddings contain diverse specific knowledge for generating the words with different part-of-speech, e.g., object/attribute embedding is good for generating nouns/adjectives. Motivated by this, we design a Mixture-of-Expert (MOE)-based decoder, where each expert is built on MHA, for discriminating the graph embeddings to generate different kinds of words. Since both the encoder and decoder are built based on the MHA, as a result, we construct a simple and homogeneous encoder-decoder unlike the previous heterogeneous ones which usually apply Fully-Connected-based GNN and LSTM-based decoder. The homogeneous architecture enables us to unify the training configuration of the whole model instead of specifying different training strategies for diverse sub-networks as in the heterogeneous pipeline, which releases the training difficulty. Extensive experiments on the MS-COCO captioning benchmark validate the effectiveness of our TFSGC. The code is in: https://anonymous.4open.science/r/ ACL23_TFSGC . * Corresponding authors. (a) Scene Graph (b) Encoder FNN Add&LN Multi-head Self Attention GNN-LSTM (c) Graph Embedding (d) Decoder MHA FNN MHA MHA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding et al.AAAI 2026
- USD: NSFW Content Detection for Text-to-Image Models via Scene GraphYuyang Zhang, Kangjie Chen, Xudong Jiang, Jiahui Wen et al.USENIX Security 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 70 citations
- Composing Object Relations and Attributes for Image-Text MatchingKhoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2024 · 31 citations
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
