Leveraging Loanword Constraints for Improving Machine Translation in a Low-Resource Multilingual Context
Felermino D. M. A. Ali, Henrique Lopes Cardoso, Rui Sousa-Silva
摘要
This research investigates how to improve machine translation systems for low-resource languages by integrating loanword constraints as external linguistic knowledge. Focusing on the Portuguese-Emakhuwa language pair, which exhibits significant lexical borrowing, we address the challenge of effectively adapting loanwords during the translation process. To tackle this, we propose a novel approach that augments source sentences with loanword constraints, explicitly linking source-language loanwords to their target-language equivalents. Then, we perform supervised fine-tuning on multilingual neural machine translation models and multiple Large Language Models of different sizes. Our results demonstrate that incorporating loanword constraints leads to significant improvements in translation quality as well as in handling loanword adaptation correctly in target languages, as measured by different machine translation metrics. This approach offers a promising direction for improving machine translation performance in low-resource settings characterized by frequent lexical borrowing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Alignment-Enhanced Transformer for Constraining NMT with Pre-Specified TranslationsKai Song, Kun Wang, Heng Yu, Yue Zhang 等AAAI 2020 · 被引用 49 次
- Lexically Constrained Neural Machine Translation with Explicit Alignment GuidanceGuanhua Chen, Yun Chen, Victor O. K. LiAAAI 2021 · 被引用 29 次
- Integrating Vectorized Lexical Constraints for Neural Machine TranslationShuo Wang, Zhixing Tan, Yang LiuACL 2022 · 被引用 12 次
- Chain-of-Dictionary Prompting Elicits Translation in Large Language ModelsHongyuan Lu, Haoran Yang, Haoyang Huang, Dongdong Zhang 等EMNLP 2024 · 被引用 9 次
相关 Paper
- Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary 等ACL 2021
- Building Resources for Emakhuwa: Machine Translation and News Classification BenchmarksFelermino Dário Mário António Ali, Henrique Lopes Cardoso, Rui Sousa-SilvaEMNLP 2024
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 被引用 11 次
- Contrastive Clustering to Mine Pseudo Parallel Data for Unsupervised TranslationXuan-Phi Nguyen, Hongyu Gong, Yun Tang, Changhan Wang 等ICLR 2022 · 被引用 5 次
- Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionFei Yuan, Linjun Shou, Xuanyu Bai, Ming Gong 等ACL 2020 · 被引用 21 次
