Prompt Learning with Quaternion Networks
Boya Shi, Zhengqin Xu, Shuai Jia, Chao Ma
Abstract
Prompt learning has emerged as an effective and data-efficient technique in large Vision-Language Models (VLMs). However, when adapting VLMs to specialized domains such as remote sensing and medical imaging, domain prompt learning remains underexplored. While large-scale domain-specific foundation models can help tackle this challenge, their concentration on a single vision level makes it challenging to prompt both vision and language modalities. To overcome this, we propose to leverage domainspecific knowledge from domain-specific foundation models to transfer the robust recognition ability of VLMs from generalized to specialized domains, using quaternion networks. Specifically, the proposed method involves using domain-specific vision features from domain-specific foundation models to guide the transformation of generalized contextual embeddings from the language branch into a specialized space within the quaternion networks. Moreover, we present a hierarchical approach that generates vision prompt features by analyzing intermodal relationships between hierarchical language prompt features and domain-specific vision features. In this way, quaternion networks can effectively mine the intermodal relationships in the specific domain, facilitating domain-specific visionlanguage contrastive learning. Extensive experiments on domain-specific datasets show that our proposed method achieves new state-of-the-art results in prompt learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial FrameQinglong Cao, Xirui Li, Ding Wang, Chao Ma et al.AAAI 2026 · 5 citations
- Diff-Prompt: Diffusion-Driven Prompt Generator with Mask SupervisionWeicai Yan, Wang Lin, Zirun Guo, Ye Wang et al.ICLR 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- Domain Prompt Learning with Quaternion NetworksQinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma et al.CVPR 2024
- Domain-Controlled Prompt LearningQinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma et al.AAAI 2024 · 38 citations
- Knowledge-Enhanced Explainable Prompting for Vision-Language ModelsYequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai et al.AAAI 2026
- Enhancing Target-unspecific Tasks through a Features MatrixFangming Cui, Yonggang Zhang, Xuan Wang, Xinmei Tian et al.ICML 2025
- Fine-Grained Prompt Learning for Face Anti-SpoofingXueli Hu, Huan Liu, Haocheng Yuan, Zhiyang Fu et al.ACM MM 2024 · 9 citations
