ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
Jonas Belouadi, Steffen Eger
摘要
State-of-the-art poetry generation systems are often complex. They either consist of task-specific model pipelines, incorporate prior knowledge in the form of manually created constraints, or both. In contrast, end-to-end models would not suffer from the overhead of having to model prior knowledge and could learn the nuances of poetry from data alone, reducing the degree of human supervision required. In this work, we investigate end-to-end poetry generation conditioned on styles such as rhyme, meter, and alliteration. We identify and address lack of training data and mismatching tokenization algorithms as possible limitations of past attempts. In particular, we successfully pre-train ByGPT5, a new token-free decoder-only language model, and fine-tune it on a large custom corpus of English and German quatrains annotated with our styles. We show that ByGPT5 outperforms other models such as mT5, ByT5, GPT-2 and ChatGPT, while also being more parameter efficient and performing favorably compared to humans. In addition, we analyze its runtime performance and demonstrate that it is not prone to memorization. We make our code, models, and datasets publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 被引用 64 次
- Evaluating Diversity in Automatic Poetry GenerationYanran Chen, Hannes Gröner, Sina Zarrieß, Steffen EgerEMNLP 2024 · 被引用 3 次
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 被引用 2 次
- LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question AnsweringRan Zhang, Wei Zhao, Lieve Macken, Steffen EgerEMNLP 2025 · 被引用 2 次
- Phun-Bench: Evaluating LLMs on Phonological Understanding in ChineseXing Yue, Yongliang Shen, Weiming LuACL 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- UBAR: Towards Fully End-to-End Task-Oriented Dialog System with GPT-2Yunyi Yang, Yunhao Li, Xiaojun QuanAAAI 2021 · 被引用 217 次
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
相关 Paper
- Don't Go Far Off: An Empirical Study on Neural Poetry TranslationTuhin Chakrabarty, Arkadiy Saakyan, Smaranda MuresanEMNLP 2021 · 被引用 8 次
- Automatic Poetry Generation from Prosaic TextTim Van de CruysACL 2020 · 被引用 54 次
- What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answeringJim Maar, Denis Paperno, Callum McDougall, Neel NandaICLR 2026 · 被引用 6 次
- GrammarT5: Grammar-Integrated Pretrained Encoder-Decoder Neural Model for CodeQihao Zhu, Qingyuan Liang, Zeyu Sun, Yingfei Xiong 等ICSE 2024 · 被引用 10 次
- Acrostic Poem GenerationRajat Agarwal, Katharina KannEMNLP 2020 · 被引用 1 次
