FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
Sebastian Joseph, Lily Chen, Jan Trienes, Hannah Louisa Göke, Monika Coers, Wei Xu, Byron C. Wallace, Junyi Jessy Li
Abstract
Plain language summarization with LLMs can be useful for improving textual accessibility of technical content. But how factual are these summaries in a high-stakes domain like medicine? This paper presents FACTPICO, a factuality benchmark for plain language summarization of medical texts describing randomized controlled trials (RCTs), which are the basis of evidence-based medicine and can directly inform patient treatment. FACTPICO consists of 345 plain language summaries of RCT abstracts generated from three LLMs (i.e., GPT-4, Llama-2, and Alpaca), with fine-grained evaluation and natural language rationales from experts. We assess the factuality of critical elements of RCTs in those summaries: Populations, Interventions, Comparators, Outcomes (PICO), as well as the reported findings concerning these. We also evaluate the correctness of the extra information (e.g., explanations) added by LLMs. Using FACTPICO, we benchmark a range of existing factuality metrics, including the newly devised ones based on LLMs. We find that plain language summarization of medical evidence is still challenging, especially when balancing between simplicity and factuality, and that existing metrics correlate poorly with expert judgments on the instance level. FactPICO and our code is available at https: //github.com/lilywchen/FactPICO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25451801-2ef4-4ddf-9771-6180b764359fCited by top-tier papers3
- SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMsYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang et al.ACL 2024 · 6 citations
- How Private are Language Models in Abstractive Summarization?Anthony Hughes, Nikolaos Aletras, Ning MaEMNLP 2025
- TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical DomainBohao Chu, Meijie Li, Sameh Frihat, Chengyu Gu et al.EMNLP 2025
Builds on9
- Fine-Tuning Language Models for FactualityKatherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning et al.ICLR 2024 · 270 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 44 citations
- Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error DetectorsLiyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban et al.ACL 2023 · 38 citations
Related papers
- Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science CommunicatorsPrasoon Bajpai, Niladri Chatterjee, Subhabrata Dutta, Tanmoy ChakrabortyEMNLP 2024
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language ModelsYancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan et al.ACL 2025
- SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical SummarizationPrakamya Mishra, Zonghai Yao, Parth Vashisht, Feiyun Ouyang et al.EMNLP 2024 · 5 citations
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsHye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. WallaceEMNLP 2023 · 11 citations
