Phrase-grounded APO for Improving Chest X-ray Report Generation
Razi Mahmood, Tanveer F. Syeda-Mahmood
摘要
The deployment of automatic radiology report generator (RRG) models in clinical workflows is being hampered by the lack of factual correctness in the produced reports. Existing methods to improve the report generators use alignment approaches that require pairs of ground truth preferred and dis-preferred responses. As these are not available at inference time in clinical workflows, new alignment methods are needed to improve report quality at inference time. In this paper, we present a new phrase-grounded automatic preference optimization (APO) alignment method which offers such improvement during inference without needing additional ground truth. Specifically, the method generates surrogate ground truth preference data for alignment automatically from the RRG model response itself though fact-checking and LLM-prompted correction. We also develop a novel APO loss function that combines preference response alignment loss with phrasal grounding loss paying attention to both the description of the finding and its image location. We show that this method of alignment, on the average, improves the report quality at inference time by 30-40% across various SOTA report generators as tested on multi-institutional chest X-ray datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang 等ICLR 2024 · 被引用 316 次
- Interactive and Explainable Region-guided Radiology Report GenerationTim Tanida, Philip Müller, Georgios Kaissis, Daniel RueckertCVPR 2023
相关 Paper
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report GenerationZhuoxiao Chen, Hongyang Yu, Ying Xu, Yadan Luo 等CVPR 2026 · 被引用 3 次
- Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference LearningQin Zhou, Guoyan Liang, Qianyi Yang, Jingyuan Chen 等ACL 2026 · 被引用 1 次
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang 等AAAI 2025 · 被引用 21 次
- CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report GenerationPablo Messina, Andrés Villa, Juan Leon Alcazar, Karen Sanchez 等CVPR 2026 · 被引用 1 次
