Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning
Yuxin Xie, Dongyue Chen, Yue Zhu, Tong Jia, Shizhuo Deng
Abstract
Image captioning aims to create natural language descriptions of images. Recent advancements in image captioning have explored text-only training methods that eliminate the need for image annotations. However, these methods are prone to generate descriptions that include objects that do not actually appear in the image, but are instead drawn from the retrieval texts or hard prompts-resulting in object hallucinations. To address this issue, we propose synergistic prompting mechanism called NASCap (Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning). Our method further improves the accuracy of generated captions by designing a fusion model with mask attention that isolate integrates retrieved captions with input features. The hard prompt mixed with negative entities is designed to further improve model's robust to wrong information. Additionally, we introduce a training-free multi-granularity fusion strategy that dynamically perceive and enhance salient regions into global representation. Extensive experiments demonstrate that NASCap sets a new state-of-the art cross-domain (transferable) captioning and performs Through extensive experiments, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in image captioning compared to zero-shot captioning based on text-only training.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Negative Entity Suppression for Zero-Shot Captioning with Synthetic ImagesZimao Lu, Hui Xu, Bing Liu, Ke WangAAAI 2026
- Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningJunjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He et al.ICCV 2023 · 80 citations
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen et al.CVPR 2024 · 38 citations
- IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot CaptioningSoeun Lee, Si-Woo Kim, Taewhan Kim, Dong-Jin KimEMNLP 2024 · 2 citations
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi et al.AAAI 2025 · 3 citations
