FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization
Nan Zhang, Yusen Zhang, Wu Guo, Prasenjit Mitra, Rui Zhang
Abstract
Summaries of medical text shall be faithful by being consistent and factual with source inputs, which is an important but understudied topic for safety and efficiency in healthcare. In this paper, we investigate and improve faithfulness in summarization on a broad range of medical summarization tasks. Our investigation reveals that current summarization models often produce unfaithful outputs for medical input text. We then introduce FAMESUMM, a framework to improve faithfulness by fine-tuning pre-trained language models based on medical knowledge. FAMESUMM performs contrastive learning on designed sets of faithful and unfaithful summaries, and it incorporates medical terms and their contexts to encourage faithful generation of medical terms. We conduct comprehensive experiments on three datasets in two languages: health question and radiology report summarization datasets in English, and a patient-doctor dialogue dataset in Chinese. Results demonstrate that FAMESUMM is flexible and effective by delivering consistent improvements over mainstream language models such as BART, T5, mT5, and PEGASUS, yielding state-of-the-art performances on metrics for faithfulness and general quality. Human evaluation by doctors also shows that FAMESUMM generates more faithful outputs. Our code is available at https: //github.com/psunlpgroup/FaMeSumm .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d48a45b8-8b87-47e6-91dc-83b8cb67b25aCited by top-tier papers2
- Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text SummarizationGunjan Balde, Soumyadeep Roy, Mainack Mondal, Niloy GangulyACL 2026
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical DomainBohao Chu, Meijie Li, Sameh Frihat, Chengyu Gu et al.EMNLP 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationShuyang Cao, Lu WangEMNLP 2021 · 130 citations
- Towards Improving Faithfulness in Abstractive SummarizationXiuying Chen, Mingzhe Li, Xin Gao, Xiangliang ZhangNeurIPS 2022 · 39 citations
- Towards Understanding Consumer Healthcare Questions on the Web with Semantically Enhanced Contrastive LearningShweta Yadav, Stefan Cobeli, Cornelia CarageaWWW 2023 · 7 citations
- uMedSum: A Unified Framework for Clinical Abstractive SummarizationAishik Nagar, Yutong Liu, Andy T. Liu, Viktor Schlegel et al.ACL 2025 · 1 citation
- Sequence Level Contrastive Learning for Text SummarizationShusheng Xu, Xingxing Zhang, Yi Wu, Furu WeiAAAI 2022 · 113 citations
