QuestEval: Summarization Asks for Fact-based Evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, Patrick Gallinari
Abstract
Summarization evaluation remains an open research problem: current metrics such as ROUGE are known to be limited and to correlate poorly with human judgments. To alleviate this issue, recent work has proposed evaluation metrics which rely on question answering models to assess whether a summary contains all the relevant information in its source document. Though promising, the proposed approaches have so far failed to correlate better than ROUGE with human judgments. In this paper, we extend previous approaches and propose a unified framework, named QU E S TEV A L. In contrast to established metrics such as ROUGE or BERTScore, QU E S TEV A L does not require any groundtruth reference. Nonetheless, QU E S TEV A L substantially improves the correlation with human judgments over four evaluation dimensions (consistency, coherence, fluency, and relevance), as shown in extensive experiments. We make code and models available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27acb10e-76d1-4c9a-b9f9-d5d0bb9c4046Cited by top-tier papers59
- What You See is What You Read? Improving Text-Image Alignment EvaluationMichal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni et al.NeurIPS 2023 · 147 citations
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao et al.EMNLP 2022 · 103 citations
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal et al.EMNLP 2024 · 53 citations
- PEER: A Collaborative Language ModelTimo Schick, Jane A. Yu, Zhengbao Jiang, Fabio Petroni et al.ICLR 2023 · 44 citations
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 44 citations
Builds on6
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 67 citations
Related papers
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai et al.EMNLP 2024 · 8 citations
- MTAS: A Reference-Free Approach for Evaluating Abstractive Summarization SystemsXiaoyan Zhu, Mingyue Jiang, Xiao-Yi Zhang, Liming Nie et al.FSE 2024 · 2 citations
- Unsupervised Reference-Free Summary Quality Evaluation via Contrastive LearningHanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa et al.EMNLP 2020 · 47 citations
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu et al.EMNLP 2022 · 16 citations
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
