Controllable Image Captioning via Prompting
Ning Wang, Jiahao Xie, Jihao Wu, Mingbo Jia, Linlin Li
Abstract
Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factual or emotional view, etc. In this paper, we show that a unified model is qualified to perform well in diverse domains and freely switch among multiple styles. Such a controllable capability is achieved by embedding the prompt learning into the image captioning framework. To be specific, we design a set of prompts to fine-tune the pre-trained image captioner. These prompts allow the model to absorb stylized data from different domains for joint training, without performance degradation in each domain. Furthermore, we optimize the prompts with learnable vectors in the continuous word embedding space, avoiding the heuristic prompt engineering and meanwhile exhibiting superior performance. In the inference stage, our model is able to generate desired stylized captions by choosing the corresponding prompts. Extensive experiments verify the controllable capability of the proposed method. Notably, we achieve outstanding performance on two diverse image captioning benchmarks including COCO Karpathy split and TextCaps using a unified model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1763a5f5-faa8-449f-baf2-04c5514ce09cCited by top-tier papers10
- PreSTU: Pre-Training for Scene-Text UnderstandingJihyung Kil, Soravit Changpinyo, Xi Chen, Hexiang Hu et al.ICCV 2023 · 39 citations
- Open-Vocabulary Calibration for Fine-tuned CLIPShuoyuan Wang, Jindong Wang, Guoqing Wang, Bob Zhang et al.ICML 2024 · 17 citations
- Cycle-Consistency Learning for Captioning and GroundingNing Wang, Jiajun Deng, Mingbo JiaAAAI 2024 · 15 citations
- Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing ImagesBo Yuan, Danpei Zhao, Zhuoran Liu, Wentao Li et al.ACM MM 2024 · 4 citations
- CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language ModelsFuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang et al.ACL 2024 · 3 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng et al.ACM MM 2022 · 8 citations
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun et al.AAAI 2024 · 7 citations
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He et al.ACM MM 2024 · 3 citations
- MemCap: Memorizing Style Knowledge for Image CaptioningWentian Zhao, Xinxiao Wu, Xiaoxun ZhangAAAI 2020 · 86 citations
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 26 citations
