Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctness
Bevan Koopman, Guido Zuccon
Abstract
This paper investigates the significant impact different prompts have on the behaviour of ChatGPT when used for health information seeking. As people more and more depend on generative large language models (LLMs) like ChatGPT, it is critical to understand model behaviour under different conditions, especially for domains where incorrect answers can have serious consequences such as health. Using the TREC Misinformation dataset, we empirically evaluate ChatGPT to show not just its effectiveness but reveal that knowledge passed in the prompt can bias the model to the detriment of answer correctness. We show this occurs both for retrieve-then-generate pipelines and based on how a user phrases their question as well as the question type. This work has important implications for the development of more robust and transparent question-answering systems based on generative large language models. Prompts, raw result files and manual analysis are made publicly available at https://github.com/ ielab/drchatgpt-health_prompting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d28eb99-496c-42cc-aaf6-8192fd9b8a80Cited by top-tier papers2
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi et al.EMNLP 2025
Builds on4
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Search Engines vs. Symptom Checkers: A Comparison of their Effectiveness for Online Health AdviceSebastian Cross, Ahmed Mourad, Guido Zuccon, Bevan KoopmanWWW 2021 · 17 citations
Related papers
- Towards Better Health Conversations: The Benefits of Context-seekingRory Sayres, Yuexing Hao, Abbi Ward, Amy Wang et al.CHI 2026 · 2 citations
- Interface Matters: Exploring Human Trust in Health Information from Large Language Models via Text, Speech, and EmbodimentXin Sun, Yunjie Liu, Jos A. Bosch, Zhuying LiCSCW 2025 · 9 citations
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
- Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to MisinformationKyubeen Han, Junseo Jang, Hongjin Kim, Geunyeong Jeong et al.ACL 2025
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin et al.EMNLP 2023 · 21 citations
