MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
Bang Yang, Fenglin Liu, Xian Wu, Yaowei Wang, Xu Sun, Yuexian Zou
摘要
Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training. However, collecting and labeling large-scale datasets is time-consuming and expensive for many scenarios and languages. Therefore, sufficient labeled pairs are usually not available. To deal with the label shortage problem, we present a simple yet effective zero-shot approach Mul-tiCapCLIP that can generate visual captions for different scenarios and languages without any labeled vision-caption pairs of downstream datasets. In the training stage, MultiCapCLIP only requires text data for input. Then it conducts two main steps: 1) retrieving concept prompts that preserve the corresponding domain knowledge of new scenarios; 2) autoencoding the prompts to learn writing styles to output captions in a desired language. In the testing stage, MultiCapCLIP instead takes visual data as input directly to retrieve the concept prompts to generate the final visual descriptions. The extensive experiments on image and video captioning across four benchmarks and four languages (i.e., English, Chinese, German, and French) confirm the effectiveness of our approach. Compared with state-of-theart zero-shot and weakly-supervised methods, our method achieves 4.8% and 21.5% absolute improvements in terms of BLEU@4 and CIDEr metrics. Our code is available at https: //github.com/yangbang18/MultiCapCLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability AssessmentTarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra 等EMNLP 2024 · 被引用 7 次
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi 等AAAI 2025 · 被引用 3 次
- Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image CaptioningJeong Ryong Lee, Yejee Shin, Geonhui Son, Dosik HwangCVPR 2025
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
相关 Paper
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 被引用 24 次
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 等ACM MM 2023 · 被引用 5 次
