Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
Wang Zhu, Tianqi Chen, Xinyan Yu, Ching Ying Lin, Jade Law, Mazen Jizzini, Jorge J. Nieva, Ruishan Liu, Robin Jia
Abstract
Cancer patients are increasingly turning to large language models (LLMs) for medical information, making it critical to assess how well these models handle complex, personalized questions. However, current medical benchmarks focus on medical exams or consumer-searched questions and do not evaluate LLMs on real patient questions with patient details. In this paper, we first have three hematology-oncology physicians evaluate cancer-related questions drawn from real patients. While LLM responses are generally accurate, the models frequently fail to recognize or address false presuppositions in the questions, posing risks to safe medical decision-making. To study this limitation systematically, we introduce Cancer-Myth, an expert-verified adversarial dataset of 585 cancer-related questions with false presuppositions. On this benchmark, no frontier LLM---including GPT-5, Gemini-2.5-Pro, and Claude-4-Sonnet---corrects these false presuppositions more than of the time. To study mitigation strategies, we further construct a 150-question Cancer-Myth-NFP set, in which physicians confirm the absence of false presuppositions. We find typical mitigation strategies, such as adding precautionary prompts with GEPA optimization, can raise accuracy on Cancer-Myth to , but at the cost of misidentifying presuppositions in of Cancer-Myth-NFP questions and causing a relative performance drop on other medical benchmarks. These findings highlight a critical gap in the reliability of LLMs, show that prompting alone is not a reliable remedy for false presuppositions, and underscore the need for more robust safeguards in medical AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59ef1c6e-82f4-4287-955d-b2b0566facd2Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 13 citations
- SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesDeqing Fu, Ameya Godbole, Robin JiaEMNLP 2023
Related papers
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsShrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang et al.EMNLP 2025 · 8 citations
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 6 citations
- How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response DraftingParker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe et al.ACL 2026
