PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3
Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A. Smith, Jiebo Luo
摘要
Knowledge-based visual question answering (VQA) involves questions that require world knowledge beyond the image to yield the correct answer. Large language models (LMs) like GPT-3 are particularly helpful for this task because of their strong knowledge retrieval and reasoning capabilities. To enable LM to understand images, prior work uses a captioning model to convert images into text. However, when summarizing an image in a single caption sentence, which visual entities to describe are often underspecified. Generic image captions often miss visual details essential for the LM to answer visual questions correctly. To address this challenge, we propose PromptCap (Prompt-guided image Captioning), a captioning model designed to serve as a better connector between images and black-box LMs. Different from generic captions, PromptCap takes a natural-language prompt to control the visual entities to describe in the generated caption. The prompt contains a question that the caption should aid in answering. To avoid extra annotation, PromptCap is trained by examples synthesized with GPT-3 and existing datasets. We demonstrate Prompt-Cap’s effectiveness on an existing pipeline in which GPT-3 is prompted with image captions to carry out VQA. Prompt-Cap outperforms generic captions by a large margin and achieves state-of-the-art accuracy on knowledge-based VQA tasks (60.4% on OK-VQA and 59.6% on A-OKVQA). Zero-shot results on WebQA show that PromptCap generalizes well to unseen domains.1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- PromptFix: You Prompt and We Fix the PhotoYongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu 等NeurIPS 2024 · 被引用 55 次
- DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsShuchang Xu, Xiaofu Jin, Huamin Qu, Yukang YanCHI 2025 · 被引用 26 次
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang 等AAAI 2025 · 被引用 18 次
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal ParallelismZedong Liu, Shenggan Cheng, Guangming Tan, Yang You 等NeurIPS 2025 · 被引用 12 次
- FlexCap: Describe Anything in Images in Controllable DetailDebidatta Dwibedi, Vidhi Jain, Jonathan Tompson, Andrew Zisserman 等NeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringZhen Yang, Zhuo Tao, Qi Chen, Liang Li 等CVPR 2025
- TOA: Task-oriented Active VQAXiaoying Xing, Mingfu Liang, Ying WuNeurIPS 2023 · 被引用 20 次
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu 等AAAI 2022 · 被引用 517 次
- Breaking the Barrier Between Pre-training and Fine-tuning: A Hybrid Prompting Model for Knowledge-Based VQAZhongfan Sun, Yongli Hu, Qingqing Gao, Huajie Jiang 等ACM MM 2023 · 被引用 6 次
- Transferable Decoding with Visual Entities for Zero-Shot Image CaptioningJunjie Fei, Teng Wang, Jinrui Zhang, Zhenyu He 等ICCV 2023 · 被引用 80 次
