IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
Soeun Lee, Si-Woo Kim, Taewhan Kim, Dong-Jin Kim
摘要
Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To address this issue, we propose a novel approach called Image-like Retrieval, which aligns text features with visually relevant features to mitigate the modality gap. Our method further enhances the accuracy of generated captions by designing a Fusion Module that integrates retrieved captions with input features. Additionally, we introduce a Frequency-based Entity Filtering technique that significantly improves caption quality. We integrate these methods into a unified framework, which we refer to as IFCap (Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning). Through extensive experimentation, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in both image captioning and video captioning compared to zero-shot captioning based on text-only training. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SIDA: Synthetic Image Driven Zero-shot Domain AdaptationYe-Chan Kim, SeungJu Cha, Si-Woo Kim, Taewhan Kim 等ACM MM 2025 · 被引用 4 次
- Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic CompositionalityYoungtaek Oh, Jae-Won Cho, Dong-Jin Kim, In So Kweon 等EMNLP 2024 · 被引用 2 次
- One Patch to Caption Them All: A Unified Zero-Shot Captioning FrameworkLorenzo Bianchi, Giacomo Pacini, Fabio Carrara, Nicola Messina 等CVPR 2026 · 被引用 2 次
- SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video CaptioningYe-Chan Kim, SeungJu Cha, Si-Woo Kim, minju Jeon 等CVPR 2026 · 被引用 1 次
- Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene DescriptionsIoanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang 等ICCV 2019 · 被引用 631 次
相关 Paper
- Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image CaptioningYuxin Xie, Dongyue Chen, Yue Zhu, Tong Jia 等ACM MM 2025
- Negative Entity Suppression for Zero-Shot Captioning with Synthetic ImagesZimao Lu, Hui Xu, Bing Liu, Ke WangAAAI 2026
- Zero-Shot Image Captioning with Multi-type Entity RepresentationsDelong Zeng, Ying Shen, Man Lin, Zihao Yi 等AAAI 2025 · 被引用 3 次
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 被引用 23 次
