HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question Answering
Zhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui Li
摘要
Visual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answering quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pretraining tasks are generally incompatible with downstream question answering tasks, which makes the knowledge from the language model not well transferable to downstream tasks, and greatly limits their performance in few-shot scenarios; (2) Under-fitting. They generally do not integrate human priors to compensate for universal knowledge from language models, so as to fit the challenging VQA problem and generate reliable answers. To address these issues, we propose HybridPrompt, a cloze-and verify-style hybrid prompt framework with bridging language models and human priors in prompt tuning for VQA. Specifically, we first modify the input questions into the cloze-style prompts to narrow the gap between upstream pre-training tasks and downstream VQA task, which ensures that the universal knowledge in the language model can be better transferred to subsequent human prior-guided prompt tuning. Then, we imitate the cognitive process of human brain to introduce topic and sample related priors to construct a dynamically learnable prompt template for human prior-guided prompt learning. Finally, we add fixedlength learnable free-parameters to further enhance the generalizability and scalability of prompt learning in the VQA model. Experimental results verify the effectiveness of Hy-bridPrompt, showing that it achieves competitive performance against previous methods on widely-used VQAv2 dataset and obtains new state-of-the-art results. Our code is released at: https://github.com/zhizhi111/hybrid .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Generative Multi-Modal Knowledge Retrieval with Large Language ModelsXinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma 等AAAI 2024 · 被引用 35 次
- Neural Residual Diffusion Models for Deep Scalable Vision GenerationZhiyuan Ma, Liangliang Zhao, Biqing Qi, Bowen ZhouNeurIPS 2024 · 被引用 15 次
- Reverse Multi-Choice Dialogue Commonsense Inference with Graph-of-ThoughtLi Zheng, Hao Fei, Fei Li, Bobo Li 等AAAI 2024 · 被引用 13 次
- Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion KnowledgeLi Zheng, Sihang Wang, Hao Fei, Zuquan Peng 等ACL 2025 · 被引用 5 次
- DreamAlign: Dynamic Text-to-3D Optimization with Human Preference AlignmentGaofeng Liu, Zhiyuan Ma, Tao FangAAAI 2025 · 被引用 4 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
相关 Paper
- Breaking the Barrier Between Pre-training and Fine-tuning: A Hybrid Prompting Model for Knowledge-Based VQAZhongfan Sun, Yongli Hu, Qingqing Gao, Huajie Jiang 等ACM MM 2023 · 被引用 6 次
- Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIPChen Huang, Skyler Seto, Samira Abnar, David Grangier 等NeurIPS 2024 · 被引用 8 次
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 被引用 16 次
- Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Pengda Qin 等ICCV 2023 · 被引用 28 次
- PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi 等ICCV 2023 · 被引用 91 次
