Facet-Aware Evaluation for Extractive Summarization
Yuning Mao, Liyuan Liu, Qi Zhu, Xiang Ren, Jiawei Han
摘要
Commonly adopted metrics for extractive summarization focus on lexical overlap at the token level. In this paper, we present a facetaware evaluation setup for better assessment of the information coverage in extracted summaries. Specifically, we treat each sentence in the reference summary as a facet, identify the sentences in the document that express the semantics of each facet as support sentences of the facet, and automatically evaluate extractive summarization methods by comparing the indices of extracted sentences and support sentences of all the facets in the reference summary. To facilitate this new evaluation setup, we construct an extractive version of the CNN/Daily Mail dataset and perform a thorough quantitative investigation, through which we demonstrate that facet-aware evaluation manifests better correlation with human judgment than ROUGE, enables fine-grained evaluation as well as comparative analysis, and reveals valuable insights of state-of-the-art summarization methods. 1 Reference: Three people in Kansas have died from a listeria outbreak. Lexical Overlap: But they did not appear identical to listeria samples taken from patients infected in the Kansas outbreak. (ROUGE-1 F1=37.0, multiple token matches but totally different semantics) Manual Extract: Five people were infected and three died in the past year in Kansas from listeria that might be linked to blue bell creameries products, according to the CDC. (ROUGE-1 F1=36.9, semantics covered but lower ROUGE due to the presence of other details) Reference: Chelsea boss Jose Mourinho and United manager Louis van Gaal are pals. Lexical Overlap: Gary Neville believes Louis van Gaal's greatest achievement as a football manager is the making of Jose Mourinho. Manual Extract: The duo have been friends since they first worked together at Barcelona in 1997 where they enjoyed a successful relationship at the Camp Nou. (ROUGE Recall/F1=0, no lexical overlap at all)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 被引用 52 次
- Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement LearningYuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren 等EMNLP 2020 · 被引用 43 次
- CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited SupervisionYuning Mao, Ming Zhong, Jiawei HanEMNLP 2022 · 被引用 11 次
- Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text GenerationYuning Mao, Wenchang Ma, Deren Lei, Jiawei Han 等EMNLP 2021 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 被引用 2 次
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 等EMNLP 2020 · 被引用 3 次
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong 等ACL 2023 · 被引用 12 次
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao 等ACL 2023 · 被引用 50 次
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang 等ACL 2020 · 被引用 410 次
