Exploring the Boundaries of GPT-4 in Radiology
Qianchu Liu, Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Maria Wetscherek, Robert Tinn, Harshita Sharma, Fernando Pérez-García, Anton Schwaighofer, Pranav Rajpurkar, Sameer Tajdin Khanna
摘要
The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we focus on assessing the performance of GPT-4, the most capable LLM so far, on the text-based applications for radiology reports, comparing against state-of-theart (SOTA) radiology-specific models. Exploring various prompting strategies, we evaluated GPT-4 on a diverse range of common radiology tasks and we found GPT-4 either outperforms or is on par with current SOTA radiology models. With zero-shot prompting, GPT-4 already obtains substantial gains (≈ 10% absolute improvement) over radiology models in temporal sentence similarity classification (accuracy) and natural language inference (F 1 ). For tasks that require learning dataset-specific style or schema (e.g. findings summarisation), GPT-4 improves with example-based prompting and matches supervised SOTA. Our extensive error analysis with a board-certified radiologist shows GPT-4 has a sufficient level of radiology knowledge with only occasional errors in complex context that require nuanced domain knowledge. For findings summarisation, GPT-4 outputs are found to be overall comparable with existing manually-written impressions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge EnhancementChe Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah 等ICML 2024 · 被引用 83 次
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa 等CHI 2024 · 被引用 81 次
- SudoLM: Learning Access Control of Parametric Knowledge with Authorization AlignmentQin Liu, Fei Wang, Chaowei Xiao, Muhao ChenACL 2025
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- LLM-RG4: Flexible and Factual Radiology Report Generation Across Diverse Input ContextsZhuhao Wang, Yihua Sun, Zihan Li, Xuan Yang 等AAAI 2025 · 被引用 6 次
- Bootstrapping Large Language Models for Radiology Report GenerationChang Liu, Yuanhe Tian, Weidong Chen, Yan Song 等AAAI 2024 · 被引用 84 次
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
- Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report GenerationQin Zhou, Guoyan Liang, Xindi Li, Jingyuan Chen 等ICCV 2025 · 被引用 2 次
