Don't Go Far Off: An Empirical Study on Neural Poetry Translation
Tuhin Chakrabarty, Arkadiy Saakyan, Smaranda Muresan
Abstract
Despite constant improvements in machine translation quality, automatic poetry translation remains a challenging problem due to the lack of open-sourced parallel poetic corpora, and to the intrinsic complexities involved in preserving the semantics, style and figurative nature of poetry. We present an empirical investigation for poetry translation along several dimensions: 1) size and style of training data (poetic vs. non-poetic), including a zero-shot setup; 2) bilingual vs. multilingual learning; and 3) language-family-specific models vs. mixed-language-family models. To accomplish this, we contribute a parallel dataset of poetry translations for several language pairs. Our results show that multilingual fine-tuning on poetic text significantly outperforms multilingual fine-tuning on non-poetic text that is 35X larger in size, both in terms of automatic metrics (BLEU, BERTScore, COMET) and human evaluation metrics such as faithfulness (meaning and poetic style). Moreover, multilingual fine-tuning on poetic data outperforms bilingual fine-tuning on poetic data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingTuhin Chakrabarty, Vishakh Padmakumar, He HeEMNLP 2022 · 42 citations
- Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureKatherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray et al.EMNLP 2022 · 23 citations
- What is the Best Way for ChatGPT to Translate Poetry?Shanshan Wang, Derek F. Wong, Jingming Yao, Lidia S. ChaoACL 2024
Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang et al.ACL 2020 · 59 citations
- Automatic Poetry Generation from Prosaic TextTim Van de CruysACL 2020 · 54 citations
Related papers
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman et al.EMNLP 2024 · 3 citations
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceAndong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai et al.EMNLP 2025
- The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM AbilitiesDavid Stap, Eva Hasler, Bill Byrne, Christof Monz et al.ACL 2024
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language ModelsJonas Belouadi, Steffen EgerACL 2023 · 12 citations
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 2 citations
