Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning
Junhao Xiao, Zhiyu Wu, Hao Lin, Yi Chen, Yahui Liu, Xiaoran Zhao, Zixu Wang, Zejiang He
Abstract
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose CLIPGlasses, a plug-and-play framework that enhances CLIP’s ability to comprehend negated visual descriptions. CLIPGlasses adapts a dual-stage design: a Lens module disentangles negated semantics from text embeddings, and a Frame module predicts context-aware repulsion strength, which is integrated into the modified similarity computation to penalize alignment with negated semantics, thereby reducing false positive matches. Experiments show that CLIP equipped with CLIPGlasses achieves competitive in-domain performance and outperforms state-of-the-art methods in cross-domain generalization. Its superiority is especially evident under low-resource conditions, indicating stronger robustness across domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1580a26-6ab1-4f8f-92e4-c7f0ed33a85aCited by top-tier papers3
- Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation EngineeringQiming Li, Xiaocheng Feng, Yixuan Ma, Ruihan Chen et al.ACL 2026 · 4 citations
- Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search TrajectoriesPeiyang Liu, Zhirui Chen, Xi Wang, Di Liang et al.ACL 2026 · 4 citations
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong et al.ACL 2026 · 3 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NERLin Sun, Jiquan Wang, Kai Zhang, Yindu Su et al.AAAI 2021 · 189 citations
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh et al.CVPR 2022 · 179 citations
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
Related papers
- Seeing What's Not There: Negation Understanding Needs More Than TrainingBhuvan Aggarwal, Amit More, Mudit Soni, Srinivasa Divakar BhatICLR 2026
- Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPJunsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu et al.ICCV 2025 · 6 citations
- Vision-Language Models Do Not Understand NegationKumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li et al.CVPR 2025
- Object-centric binding in Contrastive Language-Image PretrainingRim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal et al.NeurIPS 2025 · 14 citations
- Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsKarsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals et al.ICCV 2023 · 124 citations
