FineSurE: Fine-grained Summarization Evaluation using LLMs
Hwanjun Song, Hang Su, Igor Shalyminov, Jason Cai, Saab Mansour
摘要
Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation. Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores. This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations. To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs). It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment. We compare various open-source and proprietary LLMs as backbones for FineSurE. In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions. The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate ThemYidong Wang, Yunze Song, Tingyuan Zhu, Xuanwang Zhang 等ICLR 2026 · 被引用 29 次
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz 等ACL 2025 · 被引用 17 次
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng 等ACL 2025 · 被引用 8 次
- Disentangling Likes and Dislikes in Personalized Generative Explainable RecommendationRyotaro Shimizu, Takashi Wada, Yu Wang, Johannes Kruse 等WWW 2025 · 被引用 7 次
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau 等EMNLP 2025 · 被引用 3 次
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
相关 Paper
- One Prompt To Rule Them All: LLMs for Opinion Summary EvaluationTejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju 等ACL 2024
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMsDenis Janiak, Jakub Binkowski, Albert Sawczyn, Bogdan Gabrys 等EMNLP 2025
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme 等ACL 2024 · 被引用 27 次
- STORYSUMM: Evaluating Faithfulness in Story SummarizationMelanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams 等EMNLP 2024 · 被引用 2 次
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao 等ACL 2023 · 被引用 50 次
