Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
Piotr Sawicki, Marek Grzes, Dan Brown, Fabrício Góes
Abstract
This study adapts the Consensual Assessment Technique (CAT) for Large Language Models (LLMs), introducing a novel methodology for poetry evaluation. Using a 90-poem dataset with a ground truth based on publication venue, we demonstrate that this approach allows LLMs to significantly surpass the performance of non-expert human judges. Our method, which leverages forced-choice ranking within small, randomized batches, enabled Claude-3-Opus to achieve a Spearman's Rank Correlation of 0.87 with the ground truth, dramatically outperforming the best human nonexpert evaluation (SRC = 0.38). The LLM assessments also exhibited high inter-rater reliability, underscoring the methodology's robustness. These findings establish that LLMs, when guided by a comparative framework, can be effective and reliable tools for assessing poetry, paving the way for their broader application in other creative domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Measuring LLM Novelty As The Frontier Of Original And High-Quality OutputVishakh Padmakumar, Chen Yueh-Han, Jane Pan, Valerie Chen et al.ICLR 2026 · 6 citations
Builds on3
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Art or Artifice? Large Language Models and the False Promise of CreativityTuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan et al.CHI 2024 · 122 citations
- BatchEval: Towards Human-like Text EvaluationPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ACL 2024
Related papers
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 2 citations
- PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMsZhan Qu, Shuzhou Yuan, Michael FärberAAAI 2026 · 2 citations
- Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingTuhin Chakrabarty, Vishakh Padmakumar, He HeEMNLP 2022 · 42 citations
- SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code GenerationZhifan Ye, Jiachi Chen, Zhenzhe Shao, Lingfeng Bao et al.ASE 2025
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceAndong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai et al.EMNLP 2025
