GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
Yuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Dong Zhang, Zhehuai Chen, Engsiong Chng
Abstract
Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis selection for inference. These techniques struggle to fully exploit the rich information in the diverse N-best hypotheses, making them less optimal for translation tasks that require a single, high-quality output sequence. In this paper, we propose a new generative paradigm for translation tasks, namely GenTranslate, which builds upon LLMs to generate better results from the diverse translation versions in N-best list. Leveraging the rich linguistic knowledge and strong reasoning abilities of LLMs, our new paradigm can integrate the diverse N-best candidates to generate a higher-quality translation result. Furthermore, to support LLM finetuning, we build and release a HypoTranslate dataset that contains over 592K hypothesestranslation pairs in 11 languages. Experiments on various speech and machine translation benchmarks (e.g., FLEURS, CoVoST-2, WMT) demonstrate that our GenTranslate significantly outperforms the state-of-the-art model 1 . & MT), hold significant practical importance for 041 global communication. Similar to other NLP 042 tasks, translation tasks also gain a notable progress 043 thanks to the recent advancement of LLMs (Zhang 044 et al., 2023a; Lyu et al., 2023). In the domain of 045 speech translation, Whisper (Radford et al., 2023) 046 demonstrates superior performance by collecting 047 680k-hour data for web-scale model training. Au-048 dioPaLM2 (Rubenstein et al., 2023) integrates both 049 text-and speech-based language models into a uni-050 fied architecture to process and generate text and 051 speech, thereby augmenting speech translation per-052 formance to a great extent. On the other hand, 053 LLMs also show remarkable ability in machine 054 translation. NLLB (Costa-jussà et al., 2022) is the 055 first to extend LLMs' linguistic capability to over 056 200 languages. BigTranslate (Yang et al., 2023b) is 057 finetuned on LLaMA (Touvron et al., 2023a) with 058 multilingual instruction tuning, which achieves 059 comparable performance to ChatGPT (OpenAI, 060 2022) and Google Translate. Most recent work 061 proposes SeamlessM4T (Barrault et al., 2023a), 062 a foundational multilingual and multitask model 063 2 Related Work 113 2.1 Large Language Models 114 There is recently a surge of research interests in 115 Transformer-based large language models, such as 116 ChatGPT (OpenAI, 2022), GPT-4 (OpenAI, 2023) 117 and LLaMA (Touvron et al., 2023a,b). Benefiting 118 from the giant model size and oceans of training 119 data, LLMs can understand better the linguistic 120 structures and semantic meanings behind raw text, 121 which thus shows remarkable performance on a 122 wide range of natural language processing (NLP) 123
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbb043f9-994a-4346-91bf-0c612c306aceCited by top-tier papers4
- Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question AnsweringRunxuan Liu, Bei Luo, Jiaqi Li, Baoxin Wang et al.ACL 2025 · 21 citations
- Confident and Adaptive Generative Speech Recognition via Risk ControlAmit Damri, Bracha Laufer-GoldshteinICLR 2026
- CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language ModelsZhengdong Yang, Zhen Wan, Sheng Li, Chao-Han Huck Yang et al.EMNLP 2025
- AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question AnsweringTaolin Zhang, Dongyang Li, Chen Chen, Qizhou Chen et al.ACL 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao et al.ICML 2022 · 228 citations
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsZhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu et al.EMNLP 2023 · 200 citations
Related papers
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 122 citations
- Teaching Large Language Models to Translate with ComparisonJiali Zeng, Fandong Meng, Yongjing Yin, Jie ZhouAAAI 2024 · 74 citations
- Prompting PaLM for Translation: Assessing Strategies and PerformanceDavid Vilar, Markus Freitag, Colin Cherry, Jiaming Luo et al.ACL 2023 · 70 citations
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- LLMs Are Zero-Shot Context-Aware Simultaneous TranslatorsRoman Koshkin, Katsuhito Sudoh, Satoshi NakamuraEMNLP 2024
