Vision-Language Models Do Not Understand Negation
Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li, Philip H. S. Torr, Yoon Kim, Marzyeh Ghassemi
摘要
Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain objects but not others. Despite advancements in vision-language models (VLMs) through large-scale training, their ability to comprehend negation remains underexplored. This study addresses the question: how well do current VLMs understand negation? We introduce NegBench, a new benchmark designed to evaluate negation understanding across 18 task variations and 79k examples spanning image, video, and medical datasets. The benchmark consists of two core tasks designed to evaluate negation understanding in diverse multimodal settings: Retrieval with Negation and Multiple Choice Questions with Negated Captions. Our evaluation reveals that modern VLMs struggle significantly with negation, often performing at chance level. To address these shortcomings, we explore a data-centric approach wherein we finetune CLIP models on large-scale synthetic datasets containing millions of negated captions. We show that this approach can result in a 10% increase in recall on negated queries and a 28% boost in accuracy on multiplechoice questions with negated captions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 被引用 33 次
- CountGD++: Generalized Prompting for Open-World CountingNiki Amini-Naieni, Andrew ZissermanCVPR 2026 · 被引用 14 次
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal RetrievalSiyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao 等ICLR 2026 · 被引用 13 次
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli 等NeurIPS 2025 · 被引用 11 次
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional EvaluationRang Li, Lei Li, Shuhuai Ren, Hao Tian 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal 等ICLR 2022 · 被引用 503 次
相关 Paper
- Seeing What's Not There: Negation Understanding Needs More Than TrainingBhuvan Aggarwal, Amit More, Mudit Soni, Srinivasa Divakar BhatICLR 2026
- Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPJunsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu 等ICCV 2025 · 被引用 6 次
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao 等AAAI 2026 · 被引用 2 次
- Learn to Understand Negation in Video RetrievalZiyue Wang, Aozhu Chen, Fan Hu, Xirong LiACM MM 2022 · 被引用 12 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
