Rethinking Radiology Report Generation: From Narrative Flow to Topic-Guided Findings
Sheng Cheng, Devika Subramanian
摘要
Vision-Language Models (VLMs) for radiology report generation are typically trained to mimic the narrative flow of human experts. However, we identify a potential limitation in this conventional paradigm. We hypothesize that optimizing for narrative coherence encourages models to rely on linguistic priors and inter-sentence correlations, which can weaken their grounding in direct visual evidence and lead to factual inaccuracies. To investigate this, we design a controlled experiment demonstrating that as textual context increases, a model's reliance on the input image systematically decays. We propose LLaVA-TA (Topic-guided and Anatomy-aware), a new fine-tuning framework that directly addresses this challenge by re-engineering the generation process. Instead of producing a linear narrative, LLaVA-TA decomposes the report into a set of independent, clinically-relevant topics. By training the model to generate a discrete finding for each topic conditioned on both the full image and its corresponding anatomical region, we reduce the model's reliance on narrative flow and enforce stricter visual grounding. Our experiments show that LLaVA-TA sets a new state of the art on the MIMIC-CXR dataset, significantly improving clinical accuracy on metrics like RadGraph F1 (from 29.4 to 44.0) and CheXpert F1-14 (from 39.5 to 71.5) over strong baselines. Our work demonstrates that dismantling a report's narrative structure to enforce independent, visually-grounded observations is a crucial and effective step toward building more accurate and reliable medical VLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- Automated Structured Radiology Report GenerationJean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas 等ACL 2025
- Interactive and Explainable Region-guided Radiology Report GenerationTim Tanida, Philip Müller, Georgios Kaissis, Daniel RueckertCVPR 2023
相关 Paper
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing ModalitiesNagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye YeAAAI 2026
- LLM-RG4: Flexible and Factual Radiology Report Generation Across Diverse Input ContextsZhuhao Wang, Yihua Sun, Zihan Li, Xuan Yang 等AAAI 2025 · 被引用 6 次
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation ModelsAlice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim 等CVPR 2025
- KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationZhongzhen Huang, Xiaofan Zhang, Shaoting ZhangCVPR 2023
