EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
Sen Yang, Yu Bao, Yu Lu, Jiajun Chen, Shujian Huang, Shanbo Cheng
Abstract
Large language models (LLMs) have demonstrated strong machine translation capabilities for English-centric language pairs but underperform in direct non-English (x2x) translation. This work addresses this limitation through a synthetic data generation framework that leverages models' established English-to-x (en2x) capabilities. By extending English parallel corpora into omnidirectional datasets and developing an English-referenced quality evaluation proxy, we enable effective collection of high-quality x2x training data. Combined with preference-based optimization, our method achieves significant improvement across 72 x2x directions for widely used LLMs, while generalizing to enhance en2x performance. The results demonstrate that strategic exploitation of English-centric strengths can bootstrap comprehensive multilingual translation capabilities in LLMs. We release codes, datasets, and model checkpoints at https://github.com/ NJUNLP/EAX
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 122 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?Dawei Zhu, Pinzhen Chen, Miaoran Zhang, Barry Haddow et al.EMNLP 2024 · 3 citations
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo et al.EMNLP 2025 · 2 citations
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
- Contrastive Learning for Many-to-many Multilingual Neural Machine TranslationXiao Pan, Mingxuan Wang, Liwei Wu, Lei LiACL 2021
- X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at ScaleHaoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang et al.ICLR 2025
