Bridging the Gap between Vision and Language Domains for Improved Image Captioning
Fenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang, Wei Fan, Yuexian Zou
Abstract
Image captioning has attracted extensive research interests in recent years. Due to the great disparities between vision and language, an important goal of image captioning is to link the information in visual domain to textual domain. However, many approaches conduct this process only in the decoder, making it hard to understand the images and generate captions effectively. In this paper, we propose to bridge the gap between the vision and language domains in the encoder, by enriching visual information with textual concepts, to achieve deep image understandings. To this end, we propose to explore the textual-enriched image features. Specifically, we introduce two modules, namely Textual Distilling Module and Textual Association Module. The former distills relevant textual concepts from image features, while the latter further associates extracted concepts according to their semantics. In this manner, we acquire textual-enriched image features, which provide clear textual representations of image under no explicit supervision. The proposed approach can be used as a plugin and easily embedded into a wide range of existing image captioning systems. We conduct the extensive experiments on two benchmark image captioning datasets, i.e., MSCOCO and Flickr30k. The experimental results and analysis show that, by incorporating the proposed approach, all baseline models receive consistent improvements over all metrics, with the most significant improvement up to 10% and 9%, in terms of the task-specific metrics CIDEr and SPICE, respectively. The results demonstrate that our approach is effective and generalizes well to a wide range of models for image captioning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Audio-Oriented Multimodal Machine Comprehension via Dynamic Inter- and Intra-modality AttentionZhiqi Huang, Fenglin Liu, Xian Wu, Shen Ge et al.AAAI 2021 · 26 citations
- Distance Matters in Human-Object Interaction DetectionGuangzhi Wang, Yangyang Guo, Yongkang Wong, Mohan S. KankanhalliACM MM 2022 · 18 citations
- Selective Vision-Language Subspace Projection for Few-shot CLIPXingyu Zhu, Beier Zhu, Yi Tan, Shuo Wang et al.ACM MM 2024 · 8 citations
Builds on6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 183 citations
Related papers
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 48 citations
- Relational Distant Supervision for Image Captioning without Image-Text PairsYayun Qi, Wentian Zhao, Xinxiao WuAAAI 2024 · 5 citations
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
- Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood EstimationZihao Yue, Anwen Hu, Liang Zhang, Qin JinNeurIPS 2023 · 7 citations
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan et al.ACM MM 2021 · 31 citations
