Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
Jiaxuan Li, Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
摘要
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrievalaugmented image captioning method that prompts LLMs with object names retrieved from External Visual-name memory (EVCAP). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effortlessly augment LLMs with retrieved object names by utilizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or re-training. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EV-CAP, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pre-trained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 被引用 26 次
- Understanding Retrieval Robustness for Retrieval-augmented Image CaptioningWenyan Li, Jiaang Li, Rita Ramos, Raphael Tang 等ACL 2024 · 被引用 5 次
- Tell as You Want: Customizing Image Narrative with Knowledge and ThoughtsZiwei Yao, Qian Wang, Ruiping Wang, Xilin ChenAAAI 2026 · 被引用 1 次
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingMing Hu, Kun Yuan, Yaling Shen, Feilong Tang 等ICCV 2025 · 被引用 1 次
- Engage for All: Making Ordinary Image Descriptions Appealing Again!Yuyan Chen, Yifan Jiang, Li Zhou, Jinghan Cao 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi 等ICCV 2023 · 被引用 91 次
- NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External KnowledgeDuc Minh Vo, Hong Chen, Akihiro Sugimoto, Hideki NakayamaCVPR 2022 · 被引用 21 次
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- Smallcap: Lightweight Image Captioning Prompted with Retrieval AugmentationRita Ramos, Bruno Martins, Desmond Elliott, Yova KementchedjhievaCVPR 2023
- Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training ModelKanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu 等ACM MM 2023 · 被引用 17 次
