OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, Minlie Huang
摘要
Automatic metrics are essential for developing natural language generation (NLG) models, particularly for open-ended language generation tasks such as story generation. However, existing automatic metrics are observed to correlate poorly with human evaluation. The lack of standardized benchmark datasets makes it difficult to fully evaluate the capabilities of a metric and fairly compare different metrics. Therefore, we propose OpenMEVA, a benchmark for evaluating open-ended story generation metrics. OpenMEVA provides a comprehensive test suite to assess the capabilities of metrics, including (a) the correlation with human judgments, (b) the generalization to different model outputs and datasets, (c) the ability to judge story coherence, and (d) the robustness to perturbations. To this end, Open-MEVA includes both manually annotated stories and auto-constructed test examples. We evaluate existing metrics on OpenMEVA and observe that they have poor correlation with human judgments, fail to recognize discourselevel incoherence, and lack inferential knowledge (e.g., causal order between events), the generalization ability and robustness. Our study presents insights for developing NLG models and metrics in further research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- TaleBrush: Sketching Stories with Generative Pretrained Language ModelsJohn Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee 等CHI 2022 · 被引用 202 次
- CriticEval: Evaluating Large-scale Language Model as CriticTian Lan, Wenwei Zhang, Chen Xu, Heyan Huang 等NeurIPS 2024 · 被引用 26 次
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
- Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingJie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan 等AAAI 2024 · 被引用 8 次
- Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text EvaluationSeongYeub Chu, Jong Woo Kim, Mun Yong YiCHI 2025 · 被引用 7 次
它引用的顶会 Paper8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin 等EMNLP 2020 · 被引用 73 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 被引用 49 次
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 被引用 46 次
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan 等EMNLP 2021 · 被引用 32 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- Learning to Rank Visual Stories From Human Ranking DataChi-Yang Hsu, Yun-Wei Chu, Vincent Chen, Kuan-Chieh Lo 等ACL 2022
