Progressive Tree-Structured Prototype Network for End-to-End Image Captioning
Pengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli Gao
Abstract
Studies of image captioning are shifting towards a trend of a fully end-to-end paradigm by leveraging powerful visual pre-trained models and transformer-based generation architecture for more flexible model training and faster inference speed. State-of-the-art approaches simply extract isolated concepts or attributes to assist description generation. However, such approaches do not consider the hierarchical semantic structure in the textual domain, which leads to an unpredictable mapping between visual representations and concept words. To this end, we propose a novel Progressive Tree-Structured prototype Network (dubbed PTSN), which is the first attempt to narrow down the scope of prediction words with appropriate semantics by modeling the hierarchical textual semantics. Specifically, we design a novel embedding method called treestructured prototype, producing a set of hierarchical representative embeddings which capture the hierarchical semantic structure in textual space. To utilize such tree-structured prototypes into visual cognition, we also propose a progressive aggregation module to exploit semantic relationships within the image and prototypes. By applying our PTSN to the end-to-end captioning framework, extensive experiments conducted on MSCOCO dataset show that our method achieves a new state-of-the-art performance with 144.2% (single model) and 146.5% (ensemble of 4 models) CIDEr scores on 'Karpathy' split and 141.4% (c5) and 143.9% (c40) CIDEr scores on the official online test server. Trained models and source code have been released at: https:// github.com/ NovaMind-Z/ PTSN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a655adc-2745-4ae5-bf2a-a10aa17a4a49Cited by top-tier papers2
- A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Pengpeng Zeng et al.NeurIPS 2022 · 22 citations
- Improving Image Captioning via Predicting Structured ConceptsTing Wang, Weidong Chen, Yuanhe Tian, Yan Song et al.EMNLP 2023 · 16 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai et al.ICLR 2022 · 950 citations
- Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image EncodingPengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao et al.ICCV 2021 · 384 citations
Related papers
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 183 citations
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang et al.CVPR 2022 · 95 citations
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 178 citations
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 124 citations
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang et al.CVPR 2022 · 125 citations
