Querying as Prompt: Parameter-Efficient Learning for Multimodal Language Model
Tian Liang, Jing Huang, Ming Kong, Luyuan Chen, Qiang Zhu
Abstract
Recent advancements in language models pre-trained on large-scale corpora have significantly propelled developments in the NLP domain and advanced progress in multimodal tasks. In this paper, we propose a Parameter-Efficient multimodal language model learning strategy, named QaP (Querying as Prompt). Its core innovation is a novel modality-bridging method that allows a set of modality-specific queries to be input as soft prompts into a frozen pre-trained language model. Specifically, we introduce an efficient Text-Conditioned Resampler that is easy to incorporate into the language models, which enables adaptive injection of text-related multimodal information at different levels of the model through query learning. This approach effectively bridges multimodal information to the language models while fully leveraging its token fusion and representation potential. We validated our method across four datasets in three distinct multimodal tasks. The results demonstrate that our QaP multimodal language model achieves state-of-the-art performance in various tasks with training only 4.6% parameters. Code is available at https://github.com/Rainlt/QaP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b79f15f9-2338-4cf1-8e36-2a38ffee3f28Cited by top-tier papers3
- MHBench: Demystifying Motion Hallucination in VideoLLMsMing Kong, Xianzhou Zeng, Luyuan Chen, Yadong Li et al.AAAI 2025 · 6 citations
- MoLE: Decoding by Mixture of Layer Experts Alleviates Hallucination in Large Vision-Language ModelsTian Liang, Yuetian Du, Jing Huang, Ming Kong et al.AAAI 2025 · 3 citations
- Prototype-as-Prompt: Multimodal Sentiment Prototypes Endowing Large Language Models the Capability to Perform Multimodal Sentiment AnalysisXianbing Zhao, Lan Luo, Hengyang Lu, Buzhou TangCVPR 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
Related papers
- ATTEMPT: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft PromptsAkari Asai, Mohammadreza Salehi, Matthew E. Peters, Hannaneh HajishirziEMNLP 2022 · 55 citations
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan et al.AAAI 2024 · 102 citations
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 8 citations
- eP-ALM: Efficient Perceptual Augmentation of Language ModelsMustafa Shukor, Corentin Dancette, Matthieu CordICCV 2023 · 36 citations
- UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingHaoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu et al.ICLR 2024 · 58 citations
