Learning to Collocate Neural Modules for Image Captioning
Xu Yang, Hanwang Zhang, Jianfei Cai
Abstract
We do not speak word by word from scratch; our brain quickly structures a pattern like sth do sth at someplace and then fill in the detailed description. To render existing encoder-decoder image captioners such human-like reasoning, we propose a novel framework: learning to Collocate Neural Modules (CNM), to generate the ``inner pattern'' connecting visual encoder and language decoder. Unlike the widely-used neural module networks in visual Q&A, where the language (, question) is fully observable, CNM for captioning is more challenging as the language is being generated and thus is partially observable. To this end, we make the following technical contributions for CNM training: 1) compact module design --- one for function words and three for visual content words (, noun, adjective, and verb), 2) soft module fusion and multi-step module execution, robustifying the visual reasoning in partial observation, 3) a linguistic loss for module controller being faithful to part-of-speech collocations (, adjective is before noun). Extensive experiments on the challenging MS-COCO image captioning benchmark validate the effectiveness of our CNM image captioner. In particular, CNM achieves a new state-of-the-art 127.9 CIDEr-D on Karpathy split and a single-model 126.0 c40 on the official server. CNM is also robust to few training samples, , by training only one sentence per image, CNM can halve the performance loss compared to a strong baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1a346c6-a4ec-4a85-93de-c3e0b25bfc79Cited by top-tier papers12
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He et al.ICCV 2019 · 165 citations
- Show, Deconfound and Tell: Image Captioning with Causal InferenceBing Liu, Dong Wang, Xu Yang, Yong Zhou et al.CVPR 2022 · 66 citations
- Auto-Parsing Network for Image Captioning and Visual Question AnsweringXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiICCV 2021 · 45 citations
- Transforming Visual Scene Graphs to Image CaptionsXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu et al.ACL 2023 · 22 citations
Related papers
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 124 citations
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao et al.ICCV 2019 · 25 citations
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen et al.ICCV 2019 · 107 citations
- Bridging the Gap between Vision and Language Domains for Improved Image CaptioningFenglin Liu, Xian Wu, Shen Ge, Xiaoyu Zhang et al.ACM MM 2020 · 13 citations
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
