Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
Yuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning, Curtis P. Langlotz
摘要
Neural abstractive summarization models are able to generate summaries which have high overlap with human references. However, existing models are not optimized for factual correctness, a critical metric in real-world applications. In this work, we develop a general framework where we evaluate the factual correctness of a generated summary by factchecking it automatically against its reference using an information extraction module. We further propose a training strategy which optimizes a neural summarization model with a factual correctness reward via reinforcement learning. We apply the proposed method to the summarization of radiology reports, where factual correctness is a key requirement. On two separate datasets collected from hospitals, we show via both automatic and human evaluation that the proposed approach substantially improves the factual correctness and overall quality of outputs over a competitive neural summarization system, producing radiology summaries that approach the quality of humanauthored ones. Background: radiographic examination of the chest. clinical history: 80 years of age, male ... Findings: frontal radiograph of the chest demonstrates repositioning of the right atrial lead possibly into the ivc. ... a right apical pneumothorax can be seen from the image. moderate right and small left pleural effusions continue. no pulmonary edema is observed. heart size is upper limits of normal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Fine-Tuning Language Models for FactualityKatherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning 等ICLR 2024 · 被引用 270 次
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao 等ACL 2022 · 被引用 194 次
- Contrastive Triple Extraction with Generative TransformerHongbin Ye, Ningyu Zhang, Shumin Deng, Mosha Chen 等AAAI 2021 · 被引用 146 次
- CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationShuyang Cao, Lu WangEMNLP 2021 · 被引用 130 次
- Human Evaluation and Correlation with Automatic Metrics in Consultation Note GenerationFrancesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric 等ACL 2022 · 被引用 62 次
它引用的顶会 Paper2
相关 Paper
- Factually Consistent Summarization via Reinforcement Learning with Textual Entailment FeedbackPaul Roit, Johan Ferret, Lior Shani, Roee Aharoni 等ACL 2023 · 被引用 21 次
- Phrase-grounded APO for Improving Chest X-ray Report GenerationRazi Mahmood, Tanveer F. Syeda-MahmoodCVPR 2026
- Multi-Fact Correction in Abstractive Text SummarizationYue Dong, Shuohang Wang, Zhe Gan, Yu Cheng 等EMNLP 2020 · 被引用 99 次
- Improving Factual Consistency of Abstractive Summarization via Question AnsweringFeng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng 等ACL 2021
- Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference LearningQin Zhou, Guoyan Liang, Qianyi Yang, Jingyuan Chen 等ACL 2026 · 被引用 1 次
