Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
Wang Zhu, Tianqi Chen, Xinyan Yu, Ching Ying Lin, Jade Law, Mazen Jizzini, Jorge J. Nieva, Ruishan Liu, Robin Jia
摘要
Cancer patients are increasingly turning to large language models (LLMs) for medical information, making it critical to assess how well these models handle complex, personalized questions. However, current medical benchmarks focus on medical exams or consumer-searched questions and do not evaluate LLMs on real patient questions with patient details. In this paper, we first have three hematology-oncology physicians evaluate cancer-related questions drawn from real patients. While LLM responses are generally accurate, the models frequently fail to recognize or address false presuppositions in the questions, posing risks to safe medical decision-making. To study this limitation systematically, we introduce Cancer-Myth, an expert-verified adversarial dataset of 585 cancer-related questions with false presuppositions. On this benchmark, no frontier LLM---including GPT-5, Gemini-2.5-Pro, and Claude-4-Sonnet---corrects these false presuppositions more than of the time. To study mitigation strategies, we further construct a 150-question Cancer-Myth-NFP set, in which physicians confirm the absence of false presuppositions. We find typical mitigation strategies, such as adding precautionary prompts with GEPA optimization, can raise accuracy on Cancer-Myth to , but at the cost of misidentifying presuppositions in of Cancer-Myth-NFP questions and causing a relative performance drop on other medical benchmarks. These findings highlight a critical gap in the reliability of LLMs, show that prompting alone is not a reliable remedy for false presuppositions, and underscore the need for more robust safeguards in medical AI systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 等EMNLP 2021 · 被引用 68 次
- CREPE: Open-Domain Question Answering with False PresuppositionsXinyan Yu, Sewon Min, Luke Zettlemoyer, Hannaneh HajishirziACL 2023 · 被引用 13 次
- SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative ExamplesDeqing Fu, Ameya Godbole, Robin JiaEMNLP 2023
相关 Paper
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank 等ICLR 2026 · 被引用 24 次
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsShrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang 等EMNLP 2025 · 被引用 8 次
- Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful BeliefsMyra Cheng, Robert D. Hawkins, Dan JurafskyACL 2026 · 被引用 6 次
- How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response DraftingParker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe 等ACL 2026
