This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models
Iker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios, German Rigau
摘要
Although large language models (LLMs) have apparently acquired a certain level of grammatical knowledge and the ability to make generalizations, they fail to interpret negation, a crucial step in Natural Language Processing. We try to clarify the reasons for the sub-optimal performance of LLMs understanding negation. We introduce a large semi-automatically generated dataset of circa 400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms. We have used our dataset with the largest available open LLMs in a zero-shot approach to grasp their generalization and inference capability and we have also fine-tuned some of the models to assess whether the understanding of negation can be trained. Our findings show that, while LLMs are proficient at classifying affirmative sentences, they struggle with negative sentences and lack a deep understanding of negation, often relying on superficial cues. Although finetuning the models on negative sentences improves their performance, the lack of generalization in handling negation is persistent, highlighting the ongoing challenges of LLMs regarding negation understanding and generalization. The dataset and code are publicly available: https://github.com/hitz-zentroa/ This-is-not-a-Dataset
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPJunsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu 等ICCV 2025 · 被引用 6 次
- Vision-Language Models Do Not Understand NegationKumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li 等CVPR 2025
它引用的顶会 Paper5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- An Analysis of Natural Language Inference Benchmarks through the Lens of NegationMd Mosharaf Hossain, Venelin Kovatchev, Pranoy Dutta, Tiffany Kao 等EMNLP 2020 · 被引用 61 次
- Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeJiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng 等ACL 2023 · 被引用 23 次
相关 Paper
- How Language Models Process NegationZhejian Zhou, Tianyi Zhou, Robin Jia, Jonathan MayICML 2026
- The Impact of Negated Text on Hallucination with Large Language ModelsJaehyung Seo, Hyeonseok Moon, Heuiseok LimEMNLP 2025
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume 等EMNLP 2022 · 被引用 34 次
- Semantic Inversion, Identical Replies: Revisiting Negation Blindness in Large Language ModelsJinsung Kim, Seonmin Koo, Heuiseok LimEMNLP 2025
- Leveraging Affirmative Interpretations from Negation Improves Natural Language UnderstandingMd Mosharaf Hossain, Eduardo BlancoEMNLP 2022 · 被引用 4 次
