Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
Hannah Rashkin, Elizabeth Clark, Fantine Huot, Mirella Lapata
Abstract
Can LLMs provide support to creative writers by giving meaningful writing feedback? In this paper, we explore the challenges and limitations of model-generated writing feedback by defining a new task, dataset, and evaluation frameworks. To study model performance in a controlled manner, we present a novel test set of 1,300 stories that we corrupted to intentionally introduce writing issues. We study the performance of commonly used LLMs in this task with both automatic and human evaluation metrics. Our analysis shows that current models have strong out-of-the-box behavior in many respects -- providing specific and mostly accurate writing feedback. However, models often fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ef9b121-d6ce-4e68-83c0-f1f462a41b48Cited by top-tier papers3
- DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human ReflectionYuying Tang, Xinyi Chen, Haotian Li, Xing Xie et al.CHI 2026 · 3 citations
- MeepleLM: A Virtual Playtester Simulating Diverse Subjective ExperiencesZizhen Li, Chuanhao Li, Yibin Wang, Jianwen Sun et al.ACL 2026 · 1 citation
- Iterative Dual-Model Alignment for Story EvaluationBruce Qin, Dan GoldwasserACL 2026
Builds on16
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- Art or Artifice? Large Language Models and the False Promise of CreativityTuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan et al.CHI 2024 · 122 citations
- RewriteLM: An Instruction-Tuned Large Language Model for Text RewritingLei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu et al.AAAI 2024 · 92 citations
Related papers
- LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English WritingZhengxiang Wang, Veronika Makarova, Zhi Li, Jordan Kodner et al.ACL 2025 · 5 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student RevisionsInderjeet Nair, Jiaye Tan, Xiaotian Su, Anne Gere et al.EMNLP 2024 · 2 citations
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
- STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story GenerationNader Akoury, Shufan Wang, Josh Whiting, Stephen Hood et al.EMNLP 2020 · 5 citations
