Learning to Compose Soft Prompts for Compositional Zero-Shot Learning
Nihal V. Nayak, Peilin Yu, Stephen H. Bach
摘要
We introduce compositional soft prompting (CSP), a parameter-efficient learning technique to improve the zero-shot compositionality of large-scale pretrained vision-language models (VLMs) like CLIP. We develop CSP for compositional zero-shot learning, the task of predicting unseen attribute-object compositions (e.g., old cat and young tiger). VLMs have a flexible text encoder that can represent arbitrary classes as natural language prompts but they often underperform task-specific architectures on the compositional zero-shot benchmark datasets. CSP treats the attributes and objects that define classes as learnable tokens of vocabulary. During training, the vocabulary is tuned to recognize classes that compose tokens in multiple ways (e.g., old cat and white cat). At test time, we recompose the learned attribute-object vocabulary in new combinations to recognize novel classes. We show that CSP outperforms the CLIP on benchmark datasets by an average of 10.9 percentage points on AUC. CSP also outperforms CoOp, a soft prompting method that fine-tunes the prefix context tokens, by an average of 5.8 percentage points on AUC. We perform additional experiments to show that CSP improves generalization to higher-order attribute-attribute-object compositions (e.g., old white cat) and combinations of pretrained attributes and fine-tuned objects. The code is available at https://github.com/BatsResearch/csp.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsMatthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille 等ICCV 2023 · 被引用 51 次
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu 等NeurIPS 2023 · 被引用 49 次
- DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningZhengxiang Shi, Aldo LipaniICLR 2024 · 被引用 46 次
- Fine-tuned Language Models are Continual LearnersThomas Scialom, Tuhin Chakrabarty, Smaranda MuresanEMNLP 2022 · 被引用 46 次
- Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt TuningCristina Menghini, Andrew Delworth, Stephen H. BachNeurIPS 2023 · 被引用 43 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
相关 Paper
- Few-Shot Composition Learning for Image Retrieval with Prompt TuningJunda Wu, Rui Wang, Handong Zhao, Ruiyi Zhang 等AAAI 2023 · 被引用 16 次
- COMMA: Co-articulated Multi-Modal LearningLianyu Hu, Liqing Gao, Zekang Liu, Chi-Man Pun 等AAAI 2024 · 被引用 7 次
- TPR: Topology-Preserving Reservoirs for Generalized Zero-Shot LearningHui Chen, Yanbin Liu, Yongqiang Ma, Nanning Zheng 等NeurIPS 2024 · 被引用 7 次
- Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot LearningXiaocheng Lu, Song Guo, Ziming Liu, Jingcai GuoCVPR 2023
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan 等CVPR 2023
