Evaluating Diversity in Automatic Poetry Generation
Yanran Chen, Hannes Gröner, Sina Zarrieß, Steffen Eger
摘要
Natural Language Generation (NLG), and more generally generative AI, are among the currently most impactful research fields. Creative NLG, such as automatic poetry generation, is a fascinating niche in this area. While most previous research has focused on forms of the Turing test when evaluating automatic poetry generation -can humans distinguish between automatic and human generated poetry -we evaluate the diversity of automatically generated poetry (with a focus on quatrains), by comparing distributions of generated poetry to distributions of human poetry along structural, lexical, semantic and stylistic dimensions, assessing different model types (word vs. character-level, general purpose LLMs vs. poetry-specific models), including the very recent LLaMA3-8B, and types of fine-tuning (conditioned vs. unconditioned). We find that current automatic poetry systems are considerably underdiverse along multiple dimensions -they often do not rhyme sufficiently, are semantically too uniform and even do not match the length distribution of human poetry. Our experiments reveal, however, that style-conditioning and character-level modeling clearly increases diversity across virtually all dimensions we explore. Our identified limitations may serve as the basis for more genuinely diverse future poetry generation models. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM DiversityJiayi Zhang, Simon Yu, Derek Chong, Anthony Sicilia 等ICML 2026 · 被引用 102 次
- When Style Breaks Safety: Defending LLMs Against Superficial Style AlignmentYuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Menon Suriyakumar 等ICLR 2026 · 被引用 8 次
- Latent Planning Emerges with ScaleMichael Hanna, Emmanuel AmeisenICLR 2026 · 被引用 5 次
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 被引用 2 次
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringRan Zhang, Wei Zhao, Lieve Macken, Steffen EgerEMNLP 2025 · 被引用 2 次
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle 等ICLR 2020 · 被引用 236 次
相关 Paper
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language ModelsJonas Belouadi, Steffen EgerACL 2023 · 被引用 12 次
- so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMsSriharsh Bhyravajjula, Melanie Walsh, Anna Preus, Maria AntoniakEMNLP 2025 · 被引用 1 次
- MixPoet: Diverse Poetry Generation via Learning Controllable Mixed Latent SpaceXiaoyuan Yi, Ruoyu Li, Cheng Yang, Wenhao Li 等AAAI 2020 · 被引用 41 次
- Automatic Poetry Generation from Prosaic TextTim Van de CruysACL 2020 · 被引用 54 次
- PoetryDiffusion: Towards Joint Semantic and Metrical Manipulation in Poetry GenerationZhiyuan Hu, Chumin Liu, Yue Feng, Anh Tuan Luu 等AAAI 2024 · 被引用 11 次
