GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
Yuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Dong Zhang, Zhehuai Chen, Engsiong Chng
摘要
Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis selection for inference. These techniques struggle to fully exploit the rich information in the diverse N-best hypotheses, making them less optimal for translation tasks that require a single, high-quality output sequence. In this paper, we propose a new generative paradigm for translation tasks, namely GenTranslate, which builds upon LLMs to generate better results from the diverse translation versions in N-best list. Leveraging the rich linguistic knowledge and strong reasoning abilities of LLMs, our new paradigm can integrate the diverse N-best candidates to generate a higher-quality translation result. Furthermore, to support LLM finetuning, we build and release a HypoTranslate dataset that contains over 592K hypothesestranslation pairs in 11 languages. Experiments on various speech and machine translation benchmarks (e.g., FLEURS, CoVoST-2, WMT) demonstrate that our GenTranslate significantly outperforms the state-of-the-art model 1 . & MT), hold significant practical importance for 041 global communication. Similar to other NLP 042 tasks, translation tasks also gain a notable progress 043 thanks to the recent advancement of LLMs (Zhang 044 et al., 2023a; Lyu et al., 2023). In the domain of 045 speech translation, Whisper (Radford et al., 2023) 046 demonstrates superior performance by collecting 047 680k-hour data for web-scale model training. Au-048 dioPaLM2 (Rubenstein et al., 2023) integrates both 049 text-and speech-based language models into a uni-050 fied architecture to process and generate text and 051 speech, thereby augmenting speech translation per-052 formance to a great extent. On the other hand, 053 LLMs also show remarkable ability in machine 054 translation. NLLB (Costa-jussà et al., 2022) is the 055 first to extend LLMs' linguistic capability to over 056 200 languages. BigTranslate (Yang et al., 2023b) is 057 finetuned on LLaMA (Touvron et al., 2023a) with 058 multilingual instruction tuning, which achieves 059 comparable performance to ChatGPT (OpenAI, 060 2022) and Google Translate. Most recent work 061 proposes SeamlessM4T (Barrault et al., 2023a), 062 a foundational multilingual and multitask model 063 2 Related Work 113 2.1 Large Language Models 114 There is recently a surge of research interests in 115 Transformer-based large language models, such as 116 ChatGPT (OpenAI, 2022), GPT-4 (OpenAI, 2023) 117 and LLaMA (Touvron et al., 2023a,b). Benefiting 118 from the giant model size and oceans of training 119 data, LLMs can understand better the linguistic 120 structures and semantic meanings behind raw text, 121 which thus shows remarkable performance on a 122 wide range of natural language processing (NLP) 123
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question AnsweringRunxuan Liu, Bei Luo, Jiaqi Li, Baoxin Wang 等ACL 2025 · 被引用 21 次
- Confident and Adaptive Generative Speech Recognition via Risk ControlAmit Damri, Bracha Laufer-GoldshteinICLR 2026
- CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language ModelsZhengdong Yang, Zhen Wan, Sheng Li, Chao-Han Huck Yang 等EMNLP 2025
- AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question AnsweringTaolin Zhang, Dongyang Li, Chen Chen, Qizhou Chen 等ACL 2026
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao 等ICML 2022 · 被引用 228 次
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsZhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu 等EMNLP 2023 · 被引用 200 次
相关 Paper
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 被引用 122 次
- Teaching Large Language Models to Translate with ComparisonJiali Zeng, Fandong Meng, Yongjing Yin, Jie ZhouAAAI 2024 · 被引用 74 次
- Prompting PaLM for Translation: Assessing Strategies and PerformanceDavid Vilar, Markus Freitag, Colin Cherry, Jiaming Luo 等ACL 2023 · 被引用 70 次
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- LLMs Are Zero-Shot Context-Aware Simultaneous TranslatorsRoman Koshkin, Katsuhito Sudoh, Satoshi NakamuraEMNLP 2024
