FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization
Esin Durmus, He He, Mona T. Diab
摘要
Neural abstractive summarization models are prone to generate content inconsistent with the source document, i.e. unfaithful. Existing automatic metrics do not capture such mistakes effectively. We tackle the problem of evaluating faithfulness of a generated summary given its source document. We first collected human annotations of faithfulness for outputs from numerous models on two datasets. We find that current models exhibit a trade-off between abstractiveness and faithfulness: outputs with less word overlap with the source document are more likely to be unfaithful. Next, we propose an automatic question answering (QA) based metric for faithfulness, FEQA, 1 which leverages recent advances in reading comprehension. Given questionanswer pairs generated from the summary, a QA model extracts answers from the document; non-matched answers indicate unfaithful information in the summary. Among metrics based on word overlap, embedding similarity, and learned language understanding models, our QA-based metric has significantly higher correlation with human faithfulness scores, especially on highly abstractive summaries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper79
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Factuality Enhanced Language Models for Open-Ended Text GenerationNayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary 等NeurIPS 2022 · 被引用 318 次
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg 等ICLR 2024 · 被引用 139 次
- CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive SummarizationShuyang Cao, Lu WangEMNLP 2021 · 被引用 130 次
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao 等EMNLP 2022 · 被引用 103 次
它引用的顶会 Paper3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingMargaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck 等ACL 2020 · 被引用 120 次
相关 Paper
- Towards Improving Faithfulness in Abstractive SummarizationXiuying Chen, Mingzhe Li, Xin Gao, Xiangliang ZhangNeurIPS 2022 · 被引用 39 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams 等EMNLP 2024 · 被引用 2 次
- Multi-Fact Correction in Abstractive Text SummarizationYue Dong, Shuohang Wang, Zhe Gan, Yu Cheng 等EMNLP 2020 · 被引用 99 次
- Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive SummarizationShiyue Zhang, David Wan, Mohit BansalACL 2023 · 被引用 15 次
