CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering
Zongxi Li, Yang Li, Haoran Xie, S. Joe Qin
Abstract
Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, identifying possible implicit assumptions is crucial in QA. To address this fundamental challenge, we propose Conditional Ambiguous Question-Answering (CondAmbigQA), a benchmark comprising 2,000 ambiguous queries and condition-aware evaluation metrics 1 . Our study pioneers "conditions" as explicit contextual constraints that resolve ambiguities in QA tasks through retrievalbased annotation, where retrieved Wikipedia fragments help identify possible interpretations for a given query and annotate answers accordingly. Experiments demonstrate that models considering conditions before answering improve answer accuracy by 11.75%, with an additional 7.15% gain when conditions are explicitly provided. These results highlight that apparent hallucinations may stem from inherent query ambiguity rather than purely model failure, and demonstrate the effectiveness of condition reasoning in QA, providing researchers with tools for rigorous evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91c11072-144e-49cd-9e23-823658b49a90Cited by top-tier papers3
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation via Lens of Dynamic InteractionsNan Huo, Xiaohan Xu, Jinyang Li, Per Jacobsson et al.ICLR 2026 · 10 citations
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 1 citation
- DRIFT-BENCH: Diagnosing CoopeRative Breakdowns in LLM Agents under Input Faults via Multi-Turn InteractionHan Bao, Zheyuan Zhang, PENGCHENG JING, Zhengqing Yuan et al.ICML 2026
Builds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
- Identifying and Answering Questions with False Assumptions: An Interpretable ApproachZijie Wang, Eduardo BlancoEMNLP 2025
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang et al.EMNLP 2025
- Is There No Such Thing as a Bad Question? H4R: HalluciBot for Ratiocination, Rewriting, Ranking, and RoutingWilliam Watson, Nicole Cho, Nishan SrishankarAAAI 2025 · 7 citations
