CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
Abhilasha Ravichander, Matt Gardner, Ana Marasovic
Abstract
The full power of human language-based communication cannot be realized without negation. All human languages have some form of negation. Despite this, negation remains a challenging phenomenon for current natural language understanding systems. To facilitate the future development of models that can process negation effectively, we present CONDAQA, the first English reading comprehension dataset which requires reasoning about the implications of negated statements in paragraphs. We collect paragraphs with diverse negation cues, then have crowdworkers ask questions about the implications of the negated statement in the passage. We also have workers make three kinds of edits to the passage-paraphrasing the negated statement, changing the scope of the negation, and reversing the negation-resulting in clusters of question-answer pairs that are difficult for models to answer with spurious shortcuts. CONDAQA features 14,182 questionanswer pairs with over 200 unique negation cues and is challenging for current state-ofthe-art models. The best performing model on CONDAQA (UNIFIEDQA-V2-3B) achieves only 42% on our consistency metric, well below human performance which is 81%. We release our dataset, along with fully-finetuned, few-shot, and zero-shot evaluations, to facilitate the development of future NLP methods that work on negated language. * Work undertaken while Abhilasha Ravichander and Ana Marasović were at the Allen Institute for AI. Revision Strategy Edited Passage PARAPHRASE EDIT Complement substitution Though Philby claimed publicly in January 1988 that he did not regret his decisions and that he missed nothing about England except the only things he missed about England were some friends, Colman's mustard, and Lea & Perrins Worcestershire sauce...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeJiangjie Chen, Wei Shi, Ziquan Fu, Sijie Cheng et al.ACL 2023 · 23 citations
- ExcluIR: Exclusionary Neural Information RetrievalWenhao Zhang, Mengqi Zhang, Shiguang Wu, Jiahuan Pei et al.AAAI 2025 · 10 citations
- A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasksElie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe LanglaisEMNLP 2024 · 2 citations
- Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable QuestionsHazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr et al.EMNLP 2025 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios et al.EMNLP 2023 · 8 citations
- IfQA: A Dataset for Open-domain Question Answering under Counterfactual PresuppositionsWenhao Yu, Meng Jiang, Peter Clark, Ashish SabharwalEMNLP 2023 · 6 citations
- ComLQ: Benchmarking Complex Logical Queries in Information RetrievalGanlin Xu, Zhitao Yin, Linghao Zhang, Jiaqing Liang et al.AAAI 2026
- ReCO: A Large Scale Chinese Reading Comprehension Dataset on OpinionBingning Wang, Ting Yao, Qi Zhang, Jingfang Xu et al.AAAI 2020 · 26 citations
- "I'd rather just go to bed": Understanding Indirect AnswersAnnie Louis, Dan Roth, Filip RadlinskiEMNLP 2020
