Learning to Name Classes for Vision and Language Models
Sarah Parisot, Yongxin Yang, Steven McDonagh
摘要
Large scale vision and language models can achieve impressive zero-shot recognition performance by mapping class specific text queries to image content. Two distinct challenges that remain however, are high sensitivity to the choice of handcrafted class names that define queries, and the difficulty of adaptation to new, smaller datasets. Towards addressing these problems, we propose to leverage available data to learn, for each class, an optimal word embedding as a function of the visual content. By learning new word embeddings on an otherwise frozen model, we are able to retain zero-shot capabilities for new classes, easily adapt models to new datasets, and adjust potentially erroneous, non-descriptive or ambiguous class names. We show that our solution can easily be integrated in image classification and object detection pipelines, yields significant performance gains in multiple scenarios and provides insights into model biases and labelling errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Image-free Classifier Injection for Zero-Shot ClassificationAnders Christensen, Massimiliano Mancini, A. Sophia Koepke, Ole Winther 等ICCV 2023 · 被引用 21 次
- Renovating Names in Open-Vocabulary Segmentation BenchmarksHaiwen Huang, Songyou Peng, Dan Zhang, Andreas GeigerNeurIPS 2024 · 被引用 7 次
- FLOSS: Free Lunch in Open-Vocabulary Semantic SegmentationYasser Benigmim, Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc 等ICCV 2025 · 被引用 6 次
- ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language TasksAndrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte 等CVPR 2024 · 被引用 1 次
- Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingZiliang Chen, Xin Huang, Xiaoxuan Fan, Keze Wang 等CVPR 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
相关 Paper
- Zero-Shot Text Classification with Self-TrainingAriel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz 等EMNLP 2022 · 被引用 48 次
- CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language ModelRuijiang Dong, Zesheng Ye, Jianzhong Qi, Lei Feng 等ICLR 2026
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang 等AAAI 2024 · 被引用 7 次
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger 等NeurIPS 2023 · 被引用 63 次
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image ClassificationReza Esfandiarpoor, Stephen H. BachICLR 2024 · 被引用 18 次
