Improving Image Captioning via Predicting Structured Concepts
Ting Wang, Weidong Chen, Yuanhe Tian, Yan Song, Zhendong Mao
Abstract
Having the difficulty of solving the semantic gap between images and texts for the image captioning task, conventional studies in this area paid some attention to treating semantic concepts as a bridge between the two modalities and improved captioning performance accordingly. Although promising results on concept prediction were obtained, the aforementioned studies normally ignore the relationship among concepts, which relies on not only objects in the image, but also word dependencies in the text, so that offers a considerable potential for improving the process of generating good descriptions. In this paper, we propose a structured concept predictor (SCP) to predict concepts and their structures, then we integrate them into captioning, so as to enhance the contribution of visual signals in this task via concepts and further use their relations to distinguish cross-modal semantics for better description generation. Particularly, we design weighted graph convolutional networks (W-GCN) to depict concept relations driven by word dependencies, and then learns differentiated contributions from these concepts for following decoding process. Therefore, our approach captures potential relations among concepts and discriminatively learns different concepts, so that effectively facilitates image captioning with inherited information across modalities. Extensive experiments and their results demonstrate the effectiveness of our approach as well as each proposed module in this work. Source code is available at: https://github.com/wangting0/SCP-WGCN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79233fd8-99f8-4daa-a940-97023320a2a1Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang et al.CVPR 2022 · 125 citations
- Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image CaptioningChia-Wen Kuo, Zsolt KiraCVPR 2022 · 88 citations
Related papers
- Improving Image Captioning with Better Use of CaptionZhan Shi, Xu Zhou, Xipeng Qiu, Xiaodan ZhuACL 2020 · 84 citations
- Multimodal Graph Conditioned Diffusion Model for Video CaptioningBenhui Zhang, Junyu Gao, Yuan YuanWWW 2026
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- Bridging the Gap between Vision and Language Domains for Improved Image CaptioningFenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang et al.ACM MM 2020 · 13 citations
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningXinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia XiaoACM MM 2021 · 75 citations
