IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions
Wenhao Yu, Meng Jiang, Peter Clark, Ashish Sabharwal
Abstract
Although counterfactual reasoning is a fundamental aspect of intelligence, the lack of largescale counterfactual open-domain questionanswering (QA) benchmarks makes it difficult to evaluate and improve models on this ability. To address this void, we introduce the first such dataset, named IfQA, where each question is based on a counterfactual presupposition via an "if" clause. Such questions require models to go beyond retrieving direct factual knowledge from the Web: they must identify the right information to retrieve and reason about an imagined situation that may even go against the facts built into their parameters. The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages. Empirical analysis reveals that the IfQA dataset is highly challenging for existing open-domain QA methods, including supervised retrieve-then-read pipeline methods (F1 score 44.5), as well as recent few-shot approaches such as chain-of-thought prompting with ChatGPT (F1 score 57.2). We hope the unique challenges posed by IfQA will push open-domain QA research on both retrieval and reasoning fronts, while also helping endow counterfactual reasoning abilities to today's language understanding models. The IfQA dataset can be found and downloaded at https://allenai.org/data/ifqa .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09c1f356-b5ad-4b5e-af2c-71a038fc0801Cited by top-tier papers12
- An LLM Compiler for Parallel Function CallingSehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee et al.ICML 2024 · 142 citations
- Get an A in Math: Progressive Rectification PromptingZhenyu Wu, Meng Jiang, Chao ShenAAAI 2024 · 15 citations
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through CodeAniket Vashishtha, Qirun Dai, Hongyuan Mei, Amit Sharma et al.ICLR 2026 · 11 citations
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-PlayRan Xu, Yuchen Zhuang, Zihan Dong, Ruiyu Wang et al.NeurIPS 2025 · 10 citations
- LLMs Are Prone to Fallacies in Causal InferenceNitish Joshi, Abulhair Saparov, Yixin Wang, He HeEMNLP 2024 · 10 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 317 citations
Related papers
- What If the TV was off? Examining Counterfactual Reasoning Abilities of Multi-modal Language ModelsLetian Zhang, Xiaotong Zhai, Zhongkai Zhao, Yongshuo Zong et al.CVPR 2024 · 9 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language ModelsYuefei Chen, Vivek K. Singh, Jing Ma, Ruixiang TangAAAI 2026 · 1 citation
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou et al.EMNLP 2023 · 3 citations
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
