Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
Yuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning, Curtis P. Langlotz
Abstract
Neural abstractive summarization models are able to generate summaries which have high overlap with human references. However, existing models are not optimized for factual correctness, a critical metric in real-world applications. In this work, we develop a general framework where we evaluate the factual correctness of a generated summary by factchecking it automatically against its reference using an information extraction module. We further propose a training strategy which optimizes a neural summarization model with a factual correctness reward via reinforcement learning. We apply the proposed method to the summarization of radiology reports, where factual correctness is a key requirement. On two separate datasets collected from hospitals, we show via both automatic and human evaluation that the proposed approach substantially improves the factual correctness and overall quality of outputs over a competitive neural summarization system, producing radiology summaries that approach the quality of humanauthored ones. Background: radiographic examination of the chest. clinical history: 80 years of age, male ... Findings: frontal radiograph of the chest demonstrates repositioning of the right atrial lead possibly into the ivc. ... a right apical pneumothorax can be seen from the image. moderate right and small left pleural effusions continue. no pulmonary edema is observed. heart size is upper limits of normal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48d49439-0c26-4f7c-b6fb-fa4f8cb17cecCited by top-tier papers20
- Fine-Tuning Language Models for FactualityKatherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning et al.ICLR 2024 · 270 citations
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao et al.ACL 2022 · 194 citations
- Contrastive Triple Extraction with Generative TransformerHongbin Ye, Ningyu Zhang, Shumin Deng, Mosha Chen et al.AAAI 2021 · 146 citations
- CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationShuyang Cao, Lu WangEMNLP 2021 · 130 citations
- Human Evaluation and Correlation with Automatic Metrics in Consultation Note GenerationFrancesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric et al.ACL 2022 · 62 citations
Builds on2
Related papers
- Factually Consistent Summarization via Reinforcement Learning with Textual Entailment FeedbackPaul Roit, Johan Ferret, Lior Shani, Roee Aharoni et al.ACL 2023 · 21 citations
- Phrase-grounded APO for Improving Chest X-ray Report GenerationRazi Mahmood, Tanveer F. Syeda-MahmoodCVPR 2026
- Multi-Fact Correction in Abstractive Text SummarizationYue Dong, Shuohang Wang, Zhe Gan, Yu Cheng et al.EMNLP 2020 · 99 citations
- Improving Factual Consistency of Abstractive Summarization via Question AnsweringFeng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng et al.ACL 2021
- Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference LearningQin Zhou, Guoyan Liang, Qianyi Yang, Jingyuan Chen et al.ACL 2026 · 1 citation
