Seeing What's Not There: Negation Understanding Needs More Than Training
Bhuvan Aggarwal, Amit More, Mudit Soni, Srinivasa Divakar Bhat
摘要
Understanding the negation in a sentence is an important part of compositional understanding and logic in natural language. Many practical AI applications, such as autonomous driving, include precise instruction with negations. For example, following instruction to an AI assistant ”locate a parking spot without a vehicle” requires the assistant to not confuse between presence and absence of vehicles. Al- though joint embedding-based Vision Language Models (VLMs) like CLIP have revolutionized multi-modal tasks, they struggle to interpret negation. To address this limitation, recently many works proposed to solve the problem through a data- centric approach by introducing additional datasets with hard-negative samples for both image and text data. Contrary to these approaches, we present a zero-shot approach to tackle the negation understanding problem. We probe the properties of CLIP text embeddings and show that they follow compositional arithmetic op- erations, which allow the addition or removal of semantic information directly in the embedding space. We then present a rule-based approach to extract negated text from given caption and then use it to explicitly remove corresponding se- mantic information from original embedding, improving negation understanding in VLMs. Our approach does not require expensive training process to induce negation understanding into the model, and achieves the state-of-the-art perfor- mance on popular benchmark for negation understanding. We improve baseline CLIP model performance on NegBench from 25.5% to 67.0% for MCQ and from 50.9% to 56.1% for retrieval tasks. Even NegCLIP model which is fine-tuned on negtion datasets, our approach boosts its MCQ accuracy from 54.03% to 66.22% and retrieval accuracy from 59.25% to 60.1% showing strong performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Image Segmentation Using Text and Image PromptsTimo Lüddecke, Alexander S. EckerCVPR 2022 · 被引用 457 次
相关 Paper
- Vision-Language Models Do Not Understand NegationKumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li 等CVPR 2025
- Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPJunsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu 等ICCV 2025 · 被引用 6 次
- Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-TuningJunhao Xiao, Zhiyu Wu, Hao Lin, Yi Chen 等AAAI 2026 · 被引用 4 次
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao 等AAAI 2026 · 被引用 2 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
