SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur P. Parikh
Abstract
Reliable automatic evaluation of summarization systems is challenging due to the multifaceted and subjective nature of the task. This is especially the case for languages other than English, where human evaluations are scarce. In this work, we introduce SEAHORSE, a dataset for multilingual, multifaceted summarization evaluation. SEAHORSE consists of 96K summaries with human ratings along 6 dimensions of text quality: comprehensibility, repetition, grammar, attribution, main ideas, and conciseness. SEAHORSE covers 6 languages, 9 systems (including the reference text), and 4 summarization datasets. As a result of its size and scope, SEAHORSE can serve both as a benchmark to evaluate learnt metrics, as well as a large-scale resource for training such metrics. We show that metrics trained with SEAHORSE achieve strong performance on two out-of-domain meta-evaluation benchmarks: TRUE (Honovich et al., 2022) and mFACE (Aharoni et al., 2023) . We make the SEAHORSE dataset and metrics publicly available for future research on multilingual and multifaceted summarization evaluation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98d6c783-2bb0-45a4-a6dd-2535944c7bc4Cited by top-tier papers13
- Improving Context-Aware Preference Modeling for Language ModelsSilviu Pitis, Ziang Xiao, Nicolas Le Roux, Alessandro SordoniNeurIPS 2024 · 30 citations
- Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentZhaofeng Wu, Ananth Balashankar, Yoon Kim, Jacob Eisenstein et al.EMNLP 2024 · 29 citations
- Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language ModelsAdam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra et al.NeurIPS 2024 · 27 citations
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 19 citations
- Foundational Autoraters: Taming Large Language Models for Better Automatic EvaluationTu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar et al.EMNLP 2024 · 14 citations
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
Related papers
- Towards Multi-dimensional Evaluation of LLM Summarization across Domains and LanguagesHyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng et al.ACL 2025 · 8 citations
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li et al.ACL 2023 · 23 citations
- Automated Metrics for Medical Multi-Document Summarization Disagree with Human EvaluationsLucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong et al.ACL 2023 · 12 citations
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 52 citations
- SQuALITY: Building a Long-Document Summarization Dataset the Hard WayAlex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang et al.EMNLP 2022 · 19 citations
