Hierarchical Scene Graph Encoder-Decoder for Image Paragraph Captioning
Xu Yang, Chongyang Gao, Hanwang Zhang, Jianfei Cai
Abstract
When we humans tell a long paragraph about an image, we usually first implicitly compose a mental "script'' and then comply with it to generate the paragraph. Inspired by this, we render the modern encoder-decoder based image paragraph captioning model such ability by proposing Hierarchical Scene Graph Encoder-Decoder (HSGED) for generating coherent and distinctive paragraphs. In particular, we use the image scene graph as the "script" to incorporate rich semantic knowledge and, more importantly, the hierarchical constraints into the model. Specifically, we design a sentence scene graph RNN (SSG-RNN) to generate sub-graph level topics, which constrain the word scene graph RNN (WSG-RNN) to generate the corresponding sentences. We propose irredundant attention in SSG-RNN to improve the possibility of abstracting topics from rarely described sub-graphs and inheriting attention in WSG-RNN to generate more grounded sentences with the abstracted topics, both of which give rise to more distinctive paragraphs. An efficient sentence-level loss is also proposed for encouraging the sequence of generated sentences to be similar to that of the ground-truth paragraphs. We validate HSGED on Stanford image paragraph dataset and show that it not only achieves a new state-of-the-art 36.02 CIDEr-D, but also generates more coherent and distinctive paragraphs under various metrics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 14d26291-dd38-4198-be24-c97e96e9ec43Cited by top-tier papers7
- Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal SummarizationLitian Zhang, Xiaoming Zhang, Junshu PanAAAI 2022 · 53 citations
- Imagine That! Abstract-to-Intricate Text-to-Image Synthesis with Scene Graph Hallucination DiffusionShengqiong Wu, Hao Fei, Hanwang Zhang, Tat-Seng ChuaNeurIPS 2023 · 38 citations
- Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph GenerationXingchen Li, Long Chen, Wenbo Ma, Yi Yang et al.ACM MM 2022 · 22 citations
- TD²-Net: Toward Denoising and Debiasing for Video Scene Graph GenerationXin Lin, Chong Shi, Yibing Zhan, Zuopeng Yang et al.AAAI 2024 · 8 citations
- IC3: Image Captioning by Committee ConsensusDavid Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross et al.EMNLP 2023 · 8 citations
Related papers
- Topic Scene Graph Generation by Attention Distillation from CaptionWenbin Wang, Ruiping Wang, Xilin ChenICCV 2021 · 16 citations
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 160 citations
- Object Relation Attention for Image Paragraph CaptioningLi-Chuan Yang, Chih-Yuan Yang, Jane Yung-jen HsuAAAI 2021 · 17 citations
