OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, Minlie Huang
Abstract
Automatic metrics are essential for developing natural language generation (NLG) models, particularly for open-ended language generation tasks such as story generation. However, existing automatic metrics are observed to correlate poorly with human evaluation. The lack of standardized benchmark datasets makes it difficult to fully evaluate the capabilities of a metric and fairly compare different metrics. Therefore, we propose OpenMEVA, a benchmark for evaluating open-ended story generation metrics. OpenMEVA provides a comprehensive test suite to assess the capabilities of metrics, including (a) the correlation with human judgments, (b) the generalization to different model outputs and datasets, (c) the ability to judge story coherence, and (d) the robustness to perturbations. To this end, Open-MEVA includes both manually annotated stories and auto-constructed test examples. We evaluate existing metrics on OpenMEVA and observe that they have poor correlation with human judgments, fail to recognize discourselevel incoherence, and lack inferential knowledge (e.g., causal order between events), the generalization ability and robustness. Our study presents insights for developing NLG models and metrics in further research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 164d0148-d30e-4bc6-a187-a578b2e7b8e3Cited by top-tier papers18
- TaleBrush: Sketching Stories with Generative Pretrained Language ModelsJohn Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee et al.CHI 2022 · 202 citations
- CriticEval: Evaluating Large-scale Language Model as CriticTian Lan, Wenwei Zhang, Chen Xu, Heyan Huang et al.NeurIPS 2024 · 26 citations
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu et al.EMNLP 2024 · 17 citations
- Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingJie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan et al.AAAI 2024 · 8 citations
- Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text EvaluationSeongYeub Chu, Jong Woo Kim, Mun Yong YiCHI 2025 · 7 citations
Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue SystemsLishan Huang, Zheng Ye, Jinghui Qin, Liang Lin et al.EMNLP 2020 · 73 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 49 citations
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 46 citations
- Perturbation CheckLists for Evaluating NLG Evaluation MetricsAnanya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan et al.EMNLP 2021 · 32 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Learning to Rank Visual Stories From Human Ranking DataChi-Yang Hsu, Yun-Wei Chu, Vincent Chen, Kuan-Chieh Lo et al.ACL 2022
