Comprehending and Ordering Semantics for Image Captioning
Yehao Li, Yingwei Pan, Ting Yao, Tao Mei
Abstract
Comprehending the rich semantics in an image and ordering them in linguistic order are essential to compose a visually-grounded and linguistically coherent description for image captioning. Modern techniques commonly capitalize on a pre-trained object detector/classifier to mine the semantics in an image, while leaving the inherent linguistic ordering of semantics under-exploited. In this paper, we propose a new recipe of Transformer-style structure, namely Comprehending and Ordering Semantics Networks (COS-Net), that novelly unifies an enriched semantic comprehending and a learnable semantic ordering processes into a single architecture. Technically, we initially utilize a cross-modal retrieval model to search the relevant sentences of each image, and all words in the searched sentences are taken as primary semantic cues. Next, a novel semantic comprehender is devised to filter out the irrelevant semantic words in primary semantic cues, and mean-while infer the missing relevant semantic words visually grounded in the image. After that, we feed all the screened and enriched semantic words into a semantic ranker, which learns to allocate all semantic words in linguistic order as humans. Such sequence of ordered semantic words are further integrated with visual tokens of images to trigger sentence generation. Empirical evidences show that COS-Net clearly surpasses the state-of-the-art approaches on COCO and achieves to-date the best CIDEr score of 141.1% on Karpathy test split. Source code is available at https://github.com/YehLi/xmodaler/tree/master/configs/image_caption/cosnet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d36cbf1-2b4e-4497-9737-279393e43f09Cited by top-tier papers13
- Retrieval-Augmented Multimodal Language ModelingMichihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James et al.ICML 2023 · 153 citations
- ObjectFusion: Multi-modal 3D Object Detection with Object-Centric FusionQi Cai, Yingwei Pan, Ting Yao, Chong-Wah Ngo et al.ICCV 2023 · 71 citations
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic SegmentationYun Xing, Jian Kang, Aoran Xiao, Jiahao Nie et al.NeurIPS 2023 · 29 citations
- CgT-GAN: CLIP-guided Text GAN for Image CaptioningJiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu et al.ACM MM 2023 · 26 citations
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 183 citations
Related papers
- Semantic-Conditional Diffusion Networks for Image CaptioningJianjie Luo, Yehao Li, Yingwei Pan, Ting Yao et al.CVPR 2023
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang et al.CVPR 2022 · 125 citations
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
- Hierarchical Modular Network for Video CaptioningHanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang et al.CVPR 2022 · 95 citations
- Learning to Collocate Neural Modules for Image CaptioningXu Yang, Hanwang Zhang, Jianfei CaiICCV 2019 · 84 citations
