Don't Go Far Off: An Empirical Study on Neural Poetry Translation
Tuhin Chakrabarty, Arkadiy Saakyan, Smaranda Muresan
摘要
Despite constant improvements in machine translation quality, automatic poetry translation remains a challenging problem due to the lack of open-sourced parallel poetic corpora, and to the intrinsic complexities involved in preserving the semantics, style and figurative nature of poetry. We present an empirical investigation for poetry translation along several dimensions: 1) size and style of training data (poetic vs. non-poetic), including a zero-shot setup; 2) bilingual vs. multilingual learning; and 3) language-family-specific models vs. mixed-language-family models. To accomplish this, we contribute a parallel dataset of poetry translations for several language pairs. Our results show that multilingual fine-tuning on poetic text significantly outperforms multilingual fine-tuning on non-poetic text that is 35X larger in size, both in terms of automatic metrics (BLEU, BERTScore, COMET) and human evaluation metrics such as faithfulness (meaning and poetic style). Moreover, multilingual fine-tuning on poetic data outperforms bilingual fine-tuning on poetic data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingTuhin Chakrabarty, Vishakh Padmakumar, He HeEMNLP 2022 · 被引用 42 次
- Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureKatherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray 等EMNLP 2022 · 被引用 23 次
- What is the Best Way for ChatGPT to Translate Poetry?Shanshan Wang, Derek F. Wong, Jingming Yao, Lidia S. ChaoACL 2024
它引用的顶会 Paper13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 等ACL 2020 · 被引用 59 次
- Automatic Poetry Generation from Prosaic TextTim Van de CruysACL 2020 · 被引用 54 次
相关 Paper
- MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical LanguageShun Wang, Ge Zhang, Han Wu, Tyler Loakman 等EMNLP 2024 · 被引用 3 次
- Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceAndong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai 等EMNLP 2025
- The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM AbilitiesDavid Stap, Eva Hasler, Bill Byrne, Christof Monz 等ACL 2024
- ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language ModelsJonas Belouadi, Steffen EgerACL 2023 · 被引用 12 次
- POEMetric: The Last Stanza of HumanityBingru Li, Han Wang, Hazel WilkinsonICLR 2026 · 被引用 2 次
