CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge
Linli Yao, Weijing Chen, Qin Jin
Abstract
Automatically generating textual descriptions for massive unlabeled images on the web can greatly benefit realistic web applications, e.g. multimodal retrieval and recommendation. However, existing models suffer from the problem of generating “over-generic” descriptions, such as their tendency to generate repetitive sentences with common concepts for different images. These generic descriptions fail to provide sufficient textual semantics for ever-changing web images. Inspired by the recent success of Vision-Language Pre-training (VLP) models that learn diverse image-text concept alignment during pretraining, we explore leveraging their cross-modal pre-trained knowledge to automatically enrich the textual semantics of image descriptions. With no need for additional human annotations, we propose a plug-and-play framework, i.e CapEnrich, to complement the generic image descriptions with more semantic details. Specifically, we first propose an automatic data-building strategy to get desired training sentences, based on which we then adopt prompting strategies, i.e. learnable and template prompts, to incentivize VLP models to generate more textual details. For learnable templates, we fix the whole VLP model and only tune the prompt vectors, which leads to two advantages: 1) the pre-training knowledge of VLP models can be reserved as much as possible to describe diverse visual concepts; 2) only lightweight trainable parameters are required, so it is friendly to low data resources. Extensive experiments show that our method significantly improves the descriptiveness and diversity of generated sentences for web images. The code is available at https://github.com/yaolinli/CapEnrich.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 865edc5a-765a-4946-bb52-2185ca36ddbcCited by top-tier papers4
- Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood EstimationZihao Yue, Anwen Hu, Liang Zhang, Qin JinNeurIPS 2023 · 7 citations
- Edit As You Wish: Video Caption Editing with Multi-grained User ControlLinli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou et al.ACM MM 2024 · 4 citations
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual ReconstructionYuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang et al.EMNLP 2025 · 1 citation
- Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly SearchShuyu Yang, Yaxiong Wang, Li Zhu, Zhedong ZhengICCV 2025 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training ModelKanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu et al.ACM MM 2023 · 17 citations
- Position-Guided Text Prompt for Vision-Language Pre-TrainingJinpeng Wang, Pan Zhou, Mike Zheng Shou, Shuicheng YanCVPR 2023
- GilBERT: Generative Vision-Language Pre-Training for Image-Text RetrievalWeixiang Hong, Kaixiang Ji, Jiajia Liu, Jian Wang et al.SIGIR 2021 · 37 citations
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang et al.ICCV 2023 · 6 citations
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou et al.ACM MM 2023 · 13 citations
