TCP: Textual-Based Class-Aware Prompt Tuning for Visual-Language Model
Hantao Yao, Rui Zhang, Changsheng Xu
摘要
Prompt tuning represents a valuable technique for adapting pre-trained visual-language models (VLM) to various downstream tasks. Recent advancements in CoOp-based methods propose a set of learnable domain-shared or image-conditional textual tokens to facilitate the generation of task-specific textual classifiers. However, those textual tokens have a limited generalization ability regarding unseen domains, as they cannot dynamically adjust to the distribution of testing classes. To tackle this issue, we present a novel Textual-based Class-aware Prompt tuning(TCP) that explicitly incorporates prior knowledge about classes to en-hance their discriminability, The critical concept of TCP in-volves leveraging Textual Knowledge Embedding (TKE) to map the high generalizability of class-level textual knowledge into class-aware textual tokens. By seamlessly inte-grating these class-aware prompts into the Text Encoder, a dynamic class-aware classifier is generated to enhance dis-criminability for unseen domains. During inference, TKE dynamically generates class-aware prompts related to the unseen classes. Comprehensive evaluations demonstrate that TKE serves as a plug-and-play module effortlessly combinable with existing methods. Furthermore, TCP con-sistently achieves superior performance while demanding less training time<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://github.com/htyao89/Textual-based_Class-aware_prompt_tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- DA-Ada: Learning Domain-Aware Adapter for Domain Adaptive Object DetectionHaochen Li, Rui Zhang, Hantao Yao, Xin Zhang 等NeurIPS 2024 · 被引用 20 次
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language ModelsSilin Cheng, Kai HanNeurIPS 2025 · 被引用 7 次
- Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li 等CVPR 2026 · 被引用 7 次
- Mind the Discriminability Trap in Source-Free Cross-domain Few-shot LearningZhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Visual-Language Prompt Tuning with Knowledge-Guided Context OptimizationHantao Yao, Rui Zhang, Changsheng XuCVPR 2023
- Task-Oriented Multi-Modal Mutual Learning for Vision-Language ModelsSifan Long, Zhen Zhao, Junkun Yuan, Zichang Tan 等ICCV 2023 · 被引用 1 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Federated Text-driven Prompt Generation for Vision-Language ModelsChen Qiu, Xingyu Li, Chaithanya Kumar Mummadi, Madan Ravi Ganesh 等ICLR 2024 · 被引用 33 次
- Knowledge-Aware Prompt Tuning for Generalizable Vision-Language ModelsBaoshuo Kan, Teng Wang, Wenpeng Lu, Xiantong Zhen 等ICCV 2023 · 被引用 53 次
