Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
Zimao Lu, Hui Xu, Bing Liu, Ke Wang
摘要
Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text annotations. However, although these approaches perform well within training domains, they suffer from poor crossdomain generalization, often producing hallucinated content when encountering novel visual environments. Retrievalbased methods attempt to mitigate this limitation by leveraging external knowledge, but they can paradoxically exacerbate hallucination when retrieved captions contain entities irrelevant to the inputs. We introduce the concept of negative entities-objects that appear in generated caption but are absent from the input-and propose Negative Entity Suppression (NES) to tackle this challenge. NES seamlessly integrates three stages: (1) it employs synthetic images to ensure consistent image-to-text retrieval across both training and inference; (2) it filters negative entities from retrieved content to enhance accuracy; and (3) it applies attention-level suppression using identified negative entities to further minimize the impact of hallucination-prone features. Evaluation across multiple benchmarks demonstrates that NES maintains competitive in-domain performance while improving crossdomain transfer and reducing hallucination rates, achieving new state-of-the-art results in ZIC. Our code is available at https://github.com/nidongpinyinme/NESCap .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang 等ICCV 2019 · 被引用 631 次
- ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image CaptioningTaewhan Kim, Soeun Lee, Si-Woo Kim, Dong-Jin KimAAAI 2025 · 被引用 26 次
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 被引用 23 次
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi 等AAAI 2025 · 被引用 3 次
- Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive DecodingSicong Leng, Hang Zhang, Guanzheng Chen, Xin Li 等CVPR 2024
相关 Paper
- Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image CaptioningYuxin Xie, Dongyue Chen, Yue Zhu, Tong Jia 等ACM MM 2025
- Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningJunjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He 等ICCV 2023 · 被引用 80 次
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot CaptioningSoeun Lee, Si-Woo Kim, Taewhan Kim, Dong-Jin KimEMNLP 2024 · 被引用 2 次
- SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image CaptioningSi-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee 等ACM MM 2025
