Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style
Hongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao, Liang Sun
Abstract
Image captioning is a research hotspot where encoder-decoder models combining convolutional neural network (CNN) and long short-term memory (LSTM) achieve promising results. Despite significant progress, these models generate sentences differently from human cognitive styles. Existing models often generate a complete sentence from the first word to the end, without considering the influence of the following words on the whole sentence generation. In this paper, we explore the utilization of a human-like cognitive style, i.e., building overall cognition for the image to be described and the sentence to be constructed, for enhancing computer image understanding. This paper first proposes a Mutual-aid network structure with Bidirectional LSTMs (MaBi-LSTMs) for acquiring overall contextual information. In the training process, the forward and backward LSTMs encode the succeeding and preceding words into their respective hidden states by simultaneously constructing the whole sentence in a complementary manner. In the captioning process, the LSTM implicitly utilizes the subsequent semantic information contained in its hidden states. In fact, MaBi-LSTMs can generate two sentences in forward and backward directions. To bridge the gap between cross-domain models and generate a sentence with higher quality, we further develop a cross-modal attention mechanism to retouch the two sentences by fusing their salient parts as well as the salient areas of the image. Experimental results on the Microsoft COCO dataset show that the proposed model improves the performance of encoder-decoder models and achieves state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf5a04f2-d6c6-4efc-8280-26a876cd89c7Related papers
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 36 citations
- Learning to Collocate Neural Modules for Image CaptioningXu Yang, Hanwang Zhang, Jianfei CaiICCV 2019 · 84 citations
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen et al.AAAI 2021 · 206 citations
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 124 citations
