Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
Jiaxuan Li, Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
Abstract
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrievalaugmented image captioning method that prompts LLMs with object names retrieved from External Visual-name memory (EVCAP). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effortlessly augment LLMs with retrieved object names by utilizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or re-training. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EV-CAP, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pre-trained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 26 citations
- Understanding Retrieval Robustness for Retrieval-augmented Image CaptioningWenyan Li, Jiaang Li, Rita Ramos, Raphael Tang et al.ACL 2024 · 5 citations
- Tell as You Want: Customizing Image Narrative with Knowledge and ThoughtsZiwei Yao, Qian Wang, Ruiping Wang, Xilin ChenAAAI 2026 · 1 citation
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingMing Hu, Kun Yuan, Yaling Shen, Feilong Tang et al.ICCV 2025 · 1 citation
- Engage for All: Making Ordinary Image Descriptions Appealing Again!Yuyan Chen, Yifan Jiang, Li Zhou, Jinghan Cao et al.ICCV 2025 · 1 citation
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi et al.ICCV 2023 · 91 citations
- NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External KnowledgeDuc Minh Vo, Hong Chen, Akihiro Sugimoto, Hideki NakayamaCVPR 2022 · 21 citations
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen et al.CVPR 2024 · 38 citations
- Smallcap: Lightweight Image Captioning Prompted with Retrieval AugmentationRita Ramos, Bruno Martins, Desmond Elliott, Yova KementchedjhievaCVPR 2023
- Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training ModelKanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu et al.ACM MM 2023 · 17 citations
