Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification
Reza Esfandiarpoor, Stephen H. Bach
Abstract
A promising approach for improving the performance of vision-language models like CLIP for image classification is to extend the class descriptions (i.e., prompts) with related attributes, e.g., using brown sparrow instead of sparrow. However, current zero-shot methods select a subset of attributes regardless of commonalities between the target classes, potentially providing no useful information that would have helped to distinguish between them. For instance, they may use color instead of bill shape to distinguish between sparrows and wrens, which are both brown. We propose Follow-up Differential Descriptions (FuDD), a zero-shot approach that tailors the class descriptions to each dataset and leads to additional attributes that better differentiate the target classes. FuDD first identifies the ambiguous classes for each image, and then uses a Large Language Model (LLM) to generate new class descriptions that differentiate between them. The new class descriptions resolve the initial ambiguity and help predict the correct label. In our experiments, FuDD consistently outperforms generic description ensembles and naive LLM-generated descriptions on 12 datasets. We show that differential descriptions are an effective tool to resolve class ambiguities, which otherwise significantly degrade the performance. We also show that high quality natural language class descriptions produced by FuDD result in comparable performance to few-shot adaptation methods. Code: https: //github.com/BatsResearch/fudd
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b739875-d822-4be6-812f-9176f5e10ee3Cited by top-tier papers9
- Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud SegmentationLili Wei, Congyan Lang, Ziyi Chen, Tao Wang et al.NeurIPS 2024 · 12 citations
- Pathology-Aware Prototype Evolution via LLM-Driven Semantic Disambiguation for Multicenter Diabetic Retinopathy DiagnosisChunzheng Zhu, Yangfang Lin, Jialin Shao, Jianxin Lin et al.ACM MM 2025 · 6 citations
- Does VLM Classification Benefit from LLM Description Semantics?Pingchuan Ma, Lennart Rietdorf, Dmytro Kotovenko, Vincent Tao Hu et al.AAAI 2025 · 5 citations
- If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept DescriptionsReza Esfandiarpoor, Cristina Menghini, Stephen H. BachEMNLP 2024 · 3 citations
- DiVE-k: DIFFERENTIAL VISUAL REASONING FOR FINE-GRAINED IMAGE RECOGNITIONRaja Kumar, Arka Sadhu, Ram NevatiaICLR 2026 · 1 citation
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Language-Driven Multi-Label Zero-Shot Learning with Semantic GranularityShouwen Wang, Qian Wan, Junbin Gao, Zhigang ZengICCV 2025 · 2 citations
- Rethinking the Effect of Uninformative Class Name in Prompt LearningFengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang et al.ACM MM 2024 · 1 citation
- Improved Zero-Shot Classification by Adapting VLMs with Text DescriptionsOindrila Saha, Grant Van Horn, Subhransu MajiCVPR 2024 · 26 citations
- Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsKarsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals et al.ICCV 2023 · 124 citations
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based SelectionLincan Cai, Jingxuan Kang, Shuang Li, Wenxuan Ma et al.ICML 2025
