Facet-Aware Evaluation for Extractive Summarization
Yuning Mao, Liyuan Liu, Qi Zhu, Xiang Ren, Jiawei Han
Abstract
Commonly adopted metrics for extractive summarization focus on lexical overlap at the token level. In this paper, we present a facetaware evaluation setup for better assessment of the information coverage in extracted summaries. Specifically, we treat each sentence in the reference summary as a facet, identify the sentences in the document that express the semantics of each facet as support sentences of the facet, and automatically evaluate extractive summarization methods by comparing the indices of extracted sentences and support sentences of all the facets in the reference summary. To facilitate this new evaluation setup, we construct an extractive version of the CNN/Daily Mail dataset and perform a thorough quantitative investigation, through which we demonstrate that facet-aware evaluation manifests better correlation with human judgment than ROUGE, enables fine-grained evaluation as well as comparative analysis, and reveals valuable insights of state-of-the-art summarization methods. 1 Reference: Three people in Kansas have died from a listeria outbreak. Lexical Overlap: But they did not appear identical to listeria samples taken from patients infected in the Kansas outbreak. (ROUGE-1 F1=37.0, multiple token matches but totally different semantics) Manual Extract: Five people were infected and three died in the past year in Kansas from listeria that might be linked to blue bell creameries products, according to the CDC. (ROUGE-1 F1=36.9, semantics covered but lower ROUGE due to the presence of other details) Reference: Chelsea boss Jose Mourinho and United manager Louis van Gaal are pals. Lexical Overlap: Gary Neville believes Louis van Gaal's greatest achievement as a football manager is the making of Jose Mourinho. Manual Extract: The duo have been friends since they first worked together at Barcelona in 1997 where they enjoyed a successful relationship at the Camp Nou. (ROUGE Recall/F1=0, no lexical overlap at all)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d15ac1a2-7608-4b50-a6d9-607a24d2d3f3Cited by top-tier papers4
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 52 citations
- Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement LearningYuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren et al.EMNLP 2020 · 43 citations
- CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited SupervisionYuning Mao, Ming Zhong, Jiawei HanEMNLP 2022 · 11 citations
- Extract, Denoise and Enforce: Evaluating and Improving Concept Preservation for Text-to-Text GenerationYuning Mao, Wenchang Ma, Deren Lei, Jiawei Han et al.EMNLP 2021 · 2 citations
Builds on2
Related papers
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 2 citations
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong et al.ACL 2023 · 12 citations
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao et al.ACL 2023 · 50 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
