MeaCap: Memory-Augmented Zero-shot Image Captioning
Zequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen, Bo Chen, Zhengjue Wang
摘要
Zero-shot image captioning (IC) without well-paired image-text data can be divided into two categories, training-free and text-only-training. Generally, these two types of methods realize zero-shot IC by integrating pre-trained vision-language models like CLIP for image-text similarity evaluation and a pre-trained language model (LM) for caption generation. The main difference be-tween them is whether using a textual corpus to train the LM. Though achieving attractive performance w.r.t. some metrics, existing methods often exhibit some com-mon drawbacks. Training-free methods tend to produce hallucinations, while text-only-training often lose gener-alization capability. To move forward, in this paper, we propose a novel Memory-Augmented zero-shot image Captioning framework (MeaCap). Specifically, equipped with a textual memory, we introduce a retrieve-then-filter module to get key concepts that are highly related to the image. By deploying our proposed memory-augmented visual-related fusion score in a keywords-to-sentence LM, MeaCap can generate concept-centered captions that keep high consistency with the image with fewer hallucinations and more world-knowledge. The framework of Mea-Cap achieves the state-of-the-art performance on a se-ries of zero-shot IC settings. Our code is available at https://github.com/joeyzOz/MeaCap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Historical Test-time Prompt Tuning for Vision Foundation ModelsJingyi Zhang, Jiaxing Huang, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 29 次
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 被引用 26 次
- LaVCa: LLM-assisted Visual Cortex CaptioningTakuya Matsuyama, Shinji Nishimoto, Yu TakagiICLR 2026 · 被引用 8 次
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi 等AAAI 2025 · 被引用 3 次
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 被引用 24 次
- Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only TrainingLongtian Qiu, Shan Ning, Xuming HeAAAI 2024 · 被引用 20 次
- MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual CaptioningBang Yang, Fenglin Liu, Xian Wu, Yaowei Wang 等ACL 2023 · 被引用 10 次
- Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image CaptioningYuxin Xie, Dongyue Chen, Yue Zhu, Tong Jia 等ACM MM 2025
- RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-TrainingChen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng 等CVPR 2023
