CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge
Linli Yao, Weijing Chen, Qin Jin
摘要
Automatically generating textual descriptions for massive unlabeled images on the web can greatly benefit realistic web applications, e.g. multimodal retrieval and recommendation. However, existing models suffer from the problem of generating “over-generic” descriptions, such as their tendency to generate repetitive sentences with common concepts for different images. These generic descriptions fail to provide sufficient textual semantics for ever-changing web images. Inspired by the recent success of Vision-Language Pre-training (VLP) models that learn diverse image-text concept alignment during pretraining, we explore leveraging their cross-modal pre-trained knowledge to automatically enrich the textual semantics of image descriptions. With no need for additional human annotations, we propose a plug-and-play framework, i.e CapEnrich, to complement the generic image descriptions with more semantic details. Specifically, we first propose an automatic data-building strategy to get desired training sentences, based on which we then adopt prompting strategies, i.e. learnable and template prompts, to incentivize VLP models to generate more textual details. For learnable templates, we fix the whole VLP model and only tune the prompt vectors, which leads to two advantages: 1) the pre-training knowledge of VLP models can be reserved as much as possible to describe diverse visual concepts; 2) only lightweight trainable parameters are required, so it is friendly to low data resources. Extensive experiments show that our method significantly improves the descriptiveness and diversity of generated sentences for web images. The code is available at https://github.com/yaolinli/CapEnrich.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood EstimationZihao Yue, Anwen Hu, Liang Zhang, Qin JinNeurIPS 2023 · 被引用 7 次
- Edit As You Wish: Video Caption Editing with Multi-grained User ControlLinli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou 等ACM MM 2024 · 被引用 4 次
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual ReconstructionYuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang 等EMNLP 2025 · 被引用 1 次
- Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly SearchShuyu Yang, Yaxiong Wang, Li Zhu, Zhedong ZhengICCV 2025 · 被引用 1 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami 等NeurIPS 2021 · 被引用 1,020 次
相关 Paper
- Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training ModelKanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu 等ACM MM 2023 · 被引用 17 次
- Position-Guided Text Prompt for Vision-Language Pre-TrainingJinpeng Wang, Pan Zhou, Mike Zheng Shou, Shuicheng YanCVPR 2023
- GilBERT: Generative Vision-Language Pre-Training for Image-Text RetrievalWeixiang Hong, Kaixiang Ji, Jiajia Liu, Jian Wang 等SIGIR 2021 · 被引用 37 次
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang 等ICCV 2023 · 被引用 6 次
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou 等ACM MM 2023 · 被引用 13 次
