UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
Jian Guan, Minlie Huang
摘要
Despite the success of existing referenced metrics (e.g., BLEU and MoverScore), they correlate poorly with human judgments for openended text generation including story or dialog generation because of the notorious oneto-many issue: there are many plausible outputs for the same input, which may differ substantially in literal or semantics from the limited number of given references. To alleviate this issue, we propose UNION, a learnable UNreferenced metrIc for evaluating Open-eNded story generation, which measures the quality of a generated story without any reference. Built on top of BERT, UNION is trained to distinguish human-written stories from negative samples and recover the perturbation in negative stories. We propose an approach of constructing negative samples by mimicking the errors commonly observed in existing NLG models, including repeated plots, conflicting logic, and long-range incoherence. Experiments on two story datasets demonstrate that UNION is a reliable measure for evaluating the quality of generated stories, which correlates better with human judgments and is more generalizable than existing state-of-theart metrics. * Corresponding author Leading Context Jack was at the bar. Reference By Human He noticed a phone on the floor. He was going to take it to lost and found. But it started ringing on the way. Jack answered it and returned it to the owner's friends. Sample 1 (Reasonable, B=0.29, M=0.49, U=1.00) On the way out he noticed a phone on the floor. He asked around if anybody owned it. Eventually he gave it to the bartender. They put it into their lost and found box. Sample 2 (Reasonable, B=0.14, M=0.27, U=1.00) He had a drinking problem. He kept having more beers. After a while he passed out. When he waked up, he was surprised to find that he lost over a hundred dollars. Sample 3 (Unreasonable, B=0.20, M=0.35, U=0.00) He was going to get drunk and get drunk. The bartender told him it was already time to leave. Jack started drinking. Jack wound up returning but cops came on the way home.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
- StoryER: Automatic Story Evaluation via Ranking, Rating and ReasoningHong Chen, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao 等EMNLP 2022 · 被引用 3 次
- Dependency-based Mixture Language ModelsZhixian Yang, Xiaojun WanACL 2022 · 被引用 3 次
- Learning Personalized Alignment for Evaluating Open-ended Text GenerationDanqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang 等EMNLP 2024 · 被引用 2 次
它引用的顶会 Paper4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 被引用 49 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar 等ACL 2023 · 被引用 10 次
- Learning to Rank Visual Stories From Human Ranking DataChi-Yang Hsu, Yun-Wei Chu, Vincent Chen, Kuan-Chieh Lo 等ACL 2022
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
