Play the Shannon Game with Language Models: A Human-Free Approach to Summary Evaluation
Nicholas Egan, Oleg V. Vasilyev, John Bohannon
Abstract
The goal of a summary is to concisely state the most important information in a document. With this principle in mind, we introduce new reference-free summary evaluation metrics that use a pretrained language model to estimate the information content shared between a document and its summary. These metrics are a modern take on the Shannon Game, a method for summary quality scoring proposed decades ago, where we replace human annotators with language models. We also view these metrics as an extension of BLANC, a recently proposed approach to summary quality measurement based on the performance of a language model with and without the help of a summary. Using transformer based language models, we empirically verify that our metrics achieve state-of-the-art correlation with human judgement of the summary quality dimensions of both coherence and relevance, as well as competitive correlation with human judgement of consistency and fluency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 006a1450-4642-44d7-92f4-a1d0b6195fc4Cited by top-tier papers3
- Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationMingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing et al.EMNLP 2021 · 49 citations
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language ModelQi Jia, Siyu Ren, Yizhu Liu, Kenny Q. ZhuEMNLP 2023 · 4 citations
- COSMIC: Mutual Information for Task-Agnostic Summarization EvaluationMaxime Darrin, Philippe Formont, Jackie Chi Kit Cheung, Pablo PiantanidaACL 2024
Builds on2
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
Related papers
- A Training-free and Reference-free Summarization Evaluation Metric via Centrality-weighted Relevance and Self-referenced RedundancyWang Chen, Piji Li, Irwin KingACL 2021
- QuestEval: Summarization Asks for Fact-based EvaluationThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.EMNLP 2021
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- MTAS: A Reference-Free Approach for Evaluating Abstractive Summarization SystemsXiaoyan Zhu, Mingyue Jiang, Xiao-Yi Zhang, Liming Nie et al.FSE 2024 · 2 citations
- Unsupervised Reference-Free Summary Quality Evaluation via Contrastive LearningHanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa et al.EMNLP 2020 · 47 citations
