StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning
Hong Chen, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao, Hideki Nakayama
摘要
Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference. We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tasks: Ranking, Rating and Reasoning. Given either a machine-generated or a human-written story, StoryER requires the machine to output 1) a preference score that corresponds to human preference, 2) specific ratings and their corresponding confidences and 3) comments for various aspects (e.g., opening, character-shaping). To support these tasks, we introduce a wellannotated dataset comprising (i) 100k ranked story pairs; and (ii) a set of 46k ratings and comments on various aspects of the story. We finetune Longformer-Encoder-Decoder (LED) on the collected dataset, with the encoder responsible for preference score and aspect prediction and the decoder for comment generation. Our comprehensive experiments result in a competitive benchmark for each task, showing the high correlation to human preference. In addition, we have witnessed the joint learning of the preference scores, the aspect ratings, and the comments brings gain in each single task. Our dataset and benchmarks are publicly available to advance the research of story evaluation tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
- Learning Personalized Alignment for Evaluating Open-ended Text GenerationDanqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang 等EMNLP 2024 · 被引用 2 次
- What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story EvaluationDingyi Yang, Qin JinACL 2025 · 被引用 2 次
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance GenerationXinda Wang, Zhengxu Hou, Yangshijie Zhang, Bingren Yan 等ACL 2026 · 被引用 1 次
- DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at DischargeZonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon 等EMNLP 2025
它引用的顶会 Paper5
- UNION: An Unreferenced Metric for Evaluating Open-ended Story GenerationJian Guan, Minlie HuangEMNLP 2020 · 被引用 46 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- On The Evaluation of Machine Translation SystemsTrained With Back-TranslationSergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael AuliACL 2020 · 被引用 15 次
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu 等ACL 2021
- All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong 等ACL 2021
相关 Paper
- StoryAlign: Evaluating and Training Reward Models for Story GenerationHaotian Xia, Hao Peng, Yunjia Qi, Bin Xu 等ICLR 2026 · 被引用 2 次
- Learning to Rank Visual Stories From Human Ranking DataChi-Yang Hsu, Yun-Wei Chu, Vincent Chen, Kuan-Chieh Lo 等ACL 2022
- STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story GenerationNader Akoury, Shufan Wang, Josh Whiting, Stephen Hood 等EMNLP 2020 · 被引用 5 次
- What Makes A Good Story? Designing Composite Rewards for Visual StorytellingJunjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu 等AAAI 2020 · 被引用 73 次
- Are Large Language Models Capable of Generating Human-Level Narratives?Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang 等EMNLP 2024 · 被引用 22 次
