Examining Scaling and Transfer of Language Model Architectures for Machine Translation
Biao Zhang, Behrooz Ghorbani, Ankur Bapna, Yong Cheng, Xavier Garcia, Jonathan Shen, Orhan Firat
摘要
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing. In machine translation, EncDec has long been the favoured approach, but with few studies investigating the performance of LMs. In this work, we thoroughly examine the role of several architectural design choices on the performance of LMs on bilingual, (massively) multilingual and zero-shot translation tasks, under systematic variations of data conditions and model sizes. Our results show that: (i) Different LMs have different scaling properties, where architectural differences often have a significant impact on model performance at small scales, but the performance gap narrows as the number of parameters increases, (ii) Several design choices, including causal masking and language-modeling objectives for the source sequence, have detrimental effects on translation quality, and (iii) When paired with full-visible masking for source sequences, LMs could perform on par with EncDec on supervised bilingual and multilingual translation tasks, and improve greatly on zero-shot directions by facilitating the reduction of off-target translations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 被引用 271 次
- Pre-RMSNorm and Pre-CRMSNorm Transformers: Equivalent and Efficient Pre-LN TransformersZixuan Jiang, Jiaqi Gu, Hanqing Zhu, David Z. PanNeurIPS 2023 · 被引用 44 次
- Scaling Laws for Multilingual Neural Machine TranslationPatrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag 等ICML 2023 · 被引用 37 次
- Investigating Efficiently Extending Transformers for Long Input SummarizationJason Phang, Yao Zhao, Peter J. LiuEMNLP 2023 · 被引用 30 次
- Gemstones: A Model Suite for Multi-Faceted Scaling LawsSean McLeish, John Kirchenbauer, David Yu Miller, Siddharth Singh 等NeurIPS 2025 · 被引用 20 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Improving Massively Multilingual Neural Machine Translation and Zero-Shot TranslationBiao Zhang, Philip Williams, Ivan Titov, Rico SennrichACL 2020 · 被引用 213 次
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna 等ICLR 2022 · 被引用 130 次
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 被引用 97 次
相关 Paper
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao 等ICML 2022 · 被引用 228 次
- Jointly Masked Sequence-to-Sequence Model for Non-Autoregressive Neural Machine TranslationJunliang Guo, Linli Xu, Enhong ChenACL 2020 · 被引用 56 次
- Transforming Sequence Tagging Into A Seq2Seq TaskKarthik Raman, Iftekhar Naim, Jiecao Chen, Kazuma Hashimoto 等EMNLP 2022 · 被引用 12 次
- Mirror-Generative Neural Machine TranslationZaixiang Zheng, Hao Zhou, Shujian Huang, Lei Li 等ICLR 2020 · 被引用 37 次
- Improving Zero-Shot Translation by Disentangling Positional InformationDanni Liu, Jan Niehues, James Cross, Francisco Guzmán 等ACL 2021
