TCP: Textual-Based Class-Aware Prompt Tuning for Visual-Language Model
Hantao Yao, Rui Zhang, Changsheng Xu
Abstract
Prompt tuning represents a valuable technique for adapting pre-trained visual-language models (VLM) to various downstream tasks. Recent advancements in CoOp-based methods propose a set of learnable domain-shared or image-conditional textual tokens to facilitate the generation of task-specific textual classifiers. However, those textual tokens have a limited generalization ability regarding unseen domains, as they cannot dynamically adjust to the distribution of testing classes. To tackle this issue, we present a novel Textual-based Class-aware Prompt tuning(TCP) that explicitly incorporates prior knowledge about classes to en-hance their discriminability, The critical concept of TCP in-volves leveraging Textual Knowledge Embedding (TKE) to map the high generalizability of class-level textual knowledge into class-aware textual tokens. By seamlessly inte-grating these class-aware prompts into the Text Encoder, a dynamic class-aware classifier is generated to enhance dis-criminability for unseen domains. During inference, TKE dynamically generates class-aware prompts related to the unseen classes. Comprehensive evaluations demonstrate that TKE serves as a plug-and-play module effortlessly combinable with existing methods. Furthermore, TCP con-sistently achieves superior performance while demanding less training time<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://github.com/htyao89/Textual-based_Class-aware_prompt_tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c5ef2f9-d1a0-41b9-a59f-42f5087ebef3Cited by top-tier papers47
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang et al.NeurIPS 2024 · 20 citations
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language ModelsSilin Cheng, Kai HanNeurIPS 2025 · 7 citations
- Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li et al.CVPR 2026 · 7 citations
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li et al.CVPR 2026 · 6 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Visual-Language Prompt Tuning with Knowledge-Guided Context OptimizationHantao Yao, Rui Zhang, Changsheng XuCVPR 2023
- Task-Oriented Multi-Modal Mutual Learning for Vision-Language ModelsSifan Long, Zhen Zhao, Junkun Yuan, Zichang Tan et al.ICCV 2023 · 1 citation
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Federated Text-driven Prompt Generation for Vision-Language ModelsChen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh et al.ICLR 2024 · 33 citations
- Knowledge-Aware Prompt Tuning for Generalizable Vision-Language ModelsBaoshuo Kan, Teng Wang, Wenpeng Lu, Xiantong Zhen et al.ICCV 2023 · 53 citations
