Evaluating Diversity in Automatic Poetry Generation
Yanran Chen, Hannes Gröner, Sina Zarrieß, Steffen Eger
Abstract
Natural Language Generation (NLG), and more generally generative AI, are among the currently most impactful research fields. Creative NLG, such as automatic poetry generation, is a fascinating niche in this area. While most previous research has focused on forms of the Turing test when evaluating automatic poetry generation -can humans distinguish between automatic and human generated poetry -we evaluate the diversity of automatically generated poetry (with a focus on quatrains), by comparing distributions of generated poetry to distributions of human poetry along structural, lexical, semantic and stylistic dimensions, assessing different model types (word vs. character-level, general purpose LLMs vs. poetry-specific models), including the very recent LLaMA3-8B, and types of fine-tuning (conditioned vs. unconditioned). We find that current automatic poetry systems are considerably underdiverse along multiple dimensions -they often do not rhyme sufficiently, are semantically too uniform and even do not match the length distribution of human poetry. Our experiments reveal, however, that style-conditioning and character-level modeling clearly increases diversity across virtually all dimensions we explore. Our identified limitations may serve as the basis for more genuinely diverse future poetry generation models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91e2f6a9-af63-43b9-a18e-5e1f3d7decd2Cited by top-tier papers6
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityJiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia et al.ICML 2026 · 102 citations
- When Style Breaks Safety: Defending LLMs Against Superficial Style AlignmentYuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Menon Suriyakumar et al.ICLR 2026 · 8 citations
- Latent Planning Emerges with ScaleMichael Hanna, Emmanuel AmeisenICLR 2026 · 5 citations
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 2 citations
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringRan Zhang, Wei Zhao, Lieve Macken, Steffen EgerEMNLP 2025 · 2 citations
Builds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle et al.ICLR 2020 · 236 citations
Related papers
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language ModelsJonas Belouadi, Steffen EgerACL 2023 · 12 citations
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 1 citation
- MixPoet: Diverse Poetry Generation via Learning Controllable Mixed Latent SpaceXiaoyuan Yi, Ruoyu Li, Cheng Yang, Wenhao Li et al.AAAI 2020 · 41 citations
- Automatic Poetry Generation from Prosaic TextTim Van de CruysACL 2020 · 54 citations
- PoetryDiffusion: Towards Joint Semantic and Metrical Manipulation in Poetry GenerationZhiyuan Hu, Chumin Liu, Yue Feng, Anh Tuan Luu et al.AAAI 2024 · 11 citations
