Improving Image Captioning with Better Use of Caption
Zhan Shi, Xu Zhou, Xipeng Qiu, Xiaodan Zhu
Abstract
Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics available in captions and leverage that to enhance both image representation and caption generation. Our models first construct caption-guided visual relationship graphs that introduce beneficial inductive bias using weakly supervised multi-instance learning. The representation is then enhanced with neighbouring and contextual nodes with their textual and visual features. During generation, the model further incorporates visual relationships using multi-task learning for jointly predicting word and object/predicate tag sequences. We perform extensive experiments on the MSCOCO dataset, showing that the proposed framework significantly outperforms the baselines, resulting in the state-of-the-art performance under a wide range of evaluation metrics. The code of our paper has been made publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Multi-Modality Deep Network for Extreme Learned Image CompressionXuhao Jiang, Weimin Tan, Tian Tan, Bo Yan et al.AAAI 2023 · 27 citations
- One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing ApplicationsMengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen et al.CVPR 2024 · 16 citations
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song et al.EMNLP 2023 · 16 citations
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- PC2: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal RetrievalYue Duan, Zhangxuan Gu, Zhenzhe Ying, Lei Qi et al.ACM MM 2024 · 10 citations
Related papers
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 36 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- Image Captioning with Multimodal Guidance and Search Space OptimizationYimou Guo, Yaochen Li, Jingze Liu, Jiahui Feng et al.ACM MM 2025
- Cycle-Consistency Learning for Captioning and GroundingNing Wang, Jiajun Deng, Mingbo JiaAAAI 2024 · 15 citations
