Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation
Yixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev
摘要
Human evaluation is the foundation upon which the evaluation of both summarization systems and automatic metrics rests. However, existing human evaluation studies for summarization either exhibit a low inter-annotator agreement or have insufficient scale, and an in-depth analysis of human evaluation is lacking. Therefore, we address the shortcomings of existing summarization evaluation along the following axes: (1) We propose a modified summarization salience protocol, Atomic Content Units (ACUs), which is based on fine-grained semantic units and allows for a high interannotator agreement. (2) We curate the Robust Summarization Evaluation (RoSE) benchmark, a large human evaluation dataset consisting of 22,000 summary-level annotations over 28 top-performing systems on three datasets. (3) We conduct a comparative study of four human evaluation protocols, underscoring potential confounding factors in evaluation setups. (4) We evaluate 50 automatic metrics and their variants using the collected human annotations across evaluation protocols and demonstrate how our benchmark leads to more statistically stable and significant results. The metrics we benchmarked include recent methods based on large language models (LLMs), GPTScore and G-Eval. Furthermore, our findings have important implications for evaluating LLMs, as we show that LLMs adjusted by human feedback (e.g., GPT-3.5) may overfit unconstrained human evaluation, which is affected by the annotators' prior, input-agnostic preferences, calling for more robust, targeted evaluation methods. Statistical Power -High statistical power is difficult to reach for human evaluation of similar-performing systems. §4.1 -Increasing the sample size of human evaluation effectively raises statistical power. Summary Length -Summaries from different summarization systems show a large difference in average length. §4.2 -Difference in summary length is not well-reflected by automatic evaluation metrics. -Reference-free and reference-based human evaluation results have a near-zero correlation. Evaluation -Reference-free human evaluation strongly correlates with input-agnostic, annotator preference. Protocol Comparison -Annotator's input-agnostic preference has a strong positive correlation with summary lengths. §5.2 -Annotator's input-agnostic preference does not favor reference summaries. -Compared to smaller, fine-tuned models, zero-shot large language models (e.g. GPT-3) perform better under reference-free evaluation, but worse under reference-based evaluation. Evaluating -A higher-powered human evaluation dataset can lead to a more robust automatic metric evaluation, as shown by a tighter confidence interval and higher statistical power of metric evaluation. Automatic Metrics -Automatic metric performance differs greatly under different human evaluation protocols. §6.1 & §6.2 -Automatic metrics show relatively strong system-level correlation and moderate summary-level correlation with our robust human evaluation protocol.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 被引用 96 次
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong 等NeurIPS 2024 · 被引用 63 次
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningAhmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque 等EMNLP 2023 · 被引用 48 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
相关 Paper
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai 等ACL 2024
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 等EMNLP 2020 · 被引用 3 次
- Finding a Balanced Degree of Automation for Summary EvaluationShiyue Zhang, Mohit BansalEMNLP 2021 · 被引用 16 次
- One Prompt To Rule Them All: LLMs for Opinion Summary EvaluationTejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju 等ACL 2024
- Towards Dataset-Scale and Feature-Oriented Evaluation of Text Summarization in Large Language Model PromptsSam Yu-Te Lee, Aryaman Bahukhandi, Dongyu Liu, Kwan-Liu MaIEEE VIS 2024 · 被引用 18 次
