Small Data, Big Impact: Leveraging Minimal Data for Effective Machine Translation
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán
摘要
For many languages, machine translation progress is hindered by the lack of reliable training data. Models are trained on whatever preexisting datasets may be available and then augmented with synthetic data, because it is often not economical to pay for the creation of largescale datasets. But for the case of low-resource languages, would the creation of a few thousand professionally translated sentence pairs give any benefit? In this paper, we show that it does. We describe a broad data collection effort involving around 6k professionally translated sentence pairs for each of 39 low-resource languages, which we make publicly available. We analyse the gains of models trained on this small but high-quality data, showing that it has significant impact even when larger but lower quality pre-existing corpora are used, or when data is augmented with millions of sentences through backtranslation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine TranslationHaoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan 等ICML 2024 · 被引用 447 次
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 被引用 122 次
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech TranslationKhai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui 等EMNLP 2025
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang 等ICLR 2025
它引用的顶会 Paper1
相关 Paper
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo 等EMNLP 2025 · 被引用 2 次
- Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine TranslationXabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy WayACL 2020 · 被引用 1 次
- GATITOS: Using a New Multilingual Lexicon for Low-resource Machine TranslationAlexander Jones, Isaac Caswell, Orhan Firat, Ishank SaxenaEMNLP 2023 · 被引用 3 次
- Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary 等ACL 2021
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
