Domain Prompt Learning with Quaternion Networks
Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma, Xiaokang Yang
摘要
Prompt learning has emerged as a potent and resourceefficient technique in large Vision-Language Models (VLMs). However, its application in adapting VLMs to specialized domains like remote sensing and medical imaging, termed domain prompt learning, remains relatively unexplored. Although large-scale domain-specific foundation models offer a potential solution, their focus on a singular vision level presents challenges in prompting both vision and language modalities. To address this limitation, we propose leveraging domain-specific knowledge from these foundation models to transfer the robust recognition abilities of VLMs from generalized to specialized domains, employing quaternion networks. Our method entails utilizing domain-specific vision features from domain-specific foundation models to guide the transformation of generalized contextual embeddings from the language branch into a specialized space within quaternion networks. Furthermore, we introduce a hierarchical approach that derives vision prompt features by analyzing intermodal relationships between hierarchical language prompt features and domain-specific vision features. Through this mechanism, quaternion networks can effectively explore intermodal relationships in specific domains, facilitating domain-specific vision-language contrastive learning. Extensive experiments conducted on domain-specific datasets demonstrate that our proposed method achieves new state-of-the-art results in prompt learning. Codes are available at https: //github.com/caoql98/DPLQ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 等CVPR 2024 · 被引用 93 次
- Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference TreesSijia Chen, Yibo Wang, Yi-Feng Wu, Qingguo Chen 等NeurIPS 2024 · 被引用 51 次
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song 等NeurIPS 2024 · 被引用 41 次
- Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial FrameQinglong Cao, Xirui Li, Ding Wang, Chao Ma 等AAAI 2026 · 被引用 5 次
- Advancing Prompt Learning through an External LayerFangming Cui, Xun Yang, Chao Wu, Liang Xiao 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
相关 Paper
- Prompt Learning with Quaternion NetworksBoya Shi, Zhengqin Xu, Shuai Jia, Chao MaICLR 2024 · 被引用 4 次
- Domain-Controlled Prompt LearningQinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 等AAAI 2024 · 被引用 38 次
- Knowledge-Enhanced Explainable Prompting for Vision-Language ModelsYequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai 等AAAI 2026
- Enhancing Target-unspecific Tasks through a Features MatrixFangming Cui, Yonggang Zhang, Xuan Wang, Xinmei Tian 等ICML 2025
- Fine-Grained Prompt Learning for Face Anti-SpoofingXueli Hu, Huan Liu, Haocheng Yuan, Zhiyang Fu 等ACM MM 2024 · 被引用 9 次
