Learning to Name Classes for Vision and Language Models
Sarah Parisot, Yongxin Yang, Steven McDonagh
Abstract
Large scale vision and language models can achieve impressive zero-shot recognition performance by mapping class specific text queries to image content. Two distinct challenges that remain however, are high sensitivity to the choice of handcrafted class names that define queries, and the difficulty of adaptation to new, smaller datasets. Towards addressing these problems, we propose to leverage available data to learn, for each class, an optimal word embedding as a function of the visual content. By learning new word embeddings on an otherwise frozen model, we are able to retain zero-shot capabilities for new classes, easily adapt models to new datasets, and adjust potentially erroneous, non-descriptive or ambiguous class names. We show that our solution can easily be integrated in image classification and object detection pipelines, yields significant performance gains in multiple scenarios and provides insights into model biases and labelling errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Image-free Classifier Injection for Zero-Shot ClassificationAnders Christensen, Massimiliano Mancini, A. Sophia Koepke, Ole Winther et al.ICCV 2023 · 21 citations
- Renovating Names in Open-Vocabulary Segmentation BenchmarksHaiwen Huang, Songyou Peng, Dan Zhang, Andreas GeigerNeurIPS 2024 · 7 citations
- FLOSS: Free Lunch in Open-Vocabulary Semantic SegmentationYasser Benigmim, Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc et al.ICCV 2025 · 6 citations
- ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language TasksAndrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte et al.CVPR 2024 · 1 citation
- Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingZiliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang et al.CVPR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Zero-Shot Text Classification with Self-TrainingAriel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz et al.EMNLP 2022 · 48 citations
- CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language ModelRuijiang Dong, Zesheng Ye, Jianzhong Qi, Lei Feng et al.ICLR 2026
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang et al.AAAI 2024 · 7 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image ClassificationReza Esfandiarpoor, Stephen H. BachICLR 2024 · 18 citations
