Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctness
Bevan Koopman, Guido Zuccon
摘要
This paper investigates the significant impact different prompts have on the behaviour of ChatGPT when used for health information seeking. As people more and more depend on generative large language models (LLMs) like ChatGPT, it is critical to understand model behaviour under different conditions, especially for domains where incorrect answers can have serious consequences such as health. Using the TREC Misinformation dataset, we empirically evaluate ChatGPT to show not just its effectiveness but reveal that knowledge passed in the prompt can bias the model to the detriment of answer correctness. We show this occurs both for retrieve-then-generate pipelines and based on how a user phrases their question as well as the question type. This work has important implications for the development of more robust and transparent question-answering systems based on generative large language models. Prompts, raw result files and manual analysis are made publicly available at https://github.com/ ielab/drchatgpt-health_prompting .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice 等SIGIR 2024 · 被引用 212 次
- Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi 等EMNLP 2025
它引用的顶会 Paper4
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Search Engines vs. Symptom Checkers: A Comparison of their Effectiveness for Online Health AdviceSebastian Cross, Ahmed Mourad, Guido Zuccon, Bevan KoopmanWWW 2021 · 被引用 17 次
相关 Paper
- Towards Better Health Conversations: The Benefits of Context-seekingRory Sayres, Yuexing Hao, Abbi Ward, Amy Wang 等CHI 2026 · 被引用 2 次
- Interface Matters: Exploring Human Trust in Health Information from Large Language Models via Text, Speech, and EmbodimentXin Sun, Yunjie Liu, Jos A. Bosch, Zhuying LiCSCW 2025 · 被引用 9 次
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie 等EMNLP 2023 · 被引用 114 次
- Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to MisinformationKyubeen Han, Junseo Jang, Hongjin Kim, Geunyeong Jeong 等ACL 2025
- Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge GeneratorsLiang Chen, Yang Deng, Yatao Bian, Zeyu Qin 等EMNLP 2023 · 被引用 21 次
