Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
Oindrila Saha, Grant Van Horn, Subhransu Maji
摘要
The zero-shot performance of existing vision-language models (VLMs) such as CLIP [29] is limited by the availability of large-scale, aligned image and text datasets in specific domains. In this work, we leverage two complementary sources of information-descriptions of categories generated by large language models (LLMs) and abundant, fine-grained image classification datasets-to improve the zero-shot classification performance of VLMs across fine-grained domains. On the technical side, we develop methods to train VLMs with this “bag-level” image-text super-vision. We find that simply using these attributes at test-time does not improve performance, but our training strategy, for example, on the iNaturalist [41] dataset, leads to an average improvement of 4-5% in zero-shot classification accuracy for novel categories of birds [42] and flow-ers [23]. Similar improvements are observed in domains where a subset of the categories was used to fine-tune the model. By prompting LLMs in various ways, we generate descriptions that capture visual appearance, habitat, and geographic regions and pair them with existing attributes such as the taxonomic structure of the categories. We systematically evaluate their ability to improve zero-shot categorization in natural domains. Our findings suggest that geographic priors can be just as effective and are complementary to visual appearance. Our method also outperforms prior work on prompt-based tuning of VLMs. We release the benchmark, consisting of 14 datasets at https://github.com/cvl-umass/AdaptCLIPZS, which will contribute to future research in zero-shot recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Culture in Action: Evaluating Text-to-Image Models through Social ActivitiesSina Malakouti, Boqing Gong, Adriana KovashkaICLR 2026 · 被引用 9 次
- Retrieving Counterfactuals Improves Visual In-Context LearningGuangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong ZhangCVPR 2026 · 被引用 3 次
- Free-Grained Hierarchical Visual RecognitionSeulki Park, Zilin Wang, Stella X. YuCVPR 2026 · 被引用 3 次
- Generate, Transduct, Adapt: Iterative Transduction with VLMsOindrila Saha, Logan Lawrence, Grant Van Horn, Subhransu MajiICCV 2025 · 被引用 2 次
- RealTwin: Concept Graph Representation and Grounding Framework for Reality-Preserving Digital Twin ReconstructionZisu Li, Ruohao Li, Jiawei Li, Chao Liu 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
相关 Paper
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger 等NeurIPS 2023 · 被引用 63 次
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image ClassificationReza Esfandiarpoor, Stephen H. BachICLR 2024 · 被引用 18 次
- Large Language Models are Good Prompt Learners for Low-Shot Image ClassificationZhaoheng Zheng, Jingmin Wei, Xuefeng Hu, Haidong Zhu 等CVPR 2024 · 被引用 15 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
- Learning to Prompt with Text Only Supervision for Vision-Language ModelsMuhammad Uzair Khattak, Muhammad Ferjad Naeem, Muzammal Naseer, Luc Van Gool 等AAAI 2025 · 被引用 52 次
