Data Scaling Laws in NMT: The Effect of Noise and Architecture
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang, Colin Cherry, Behnam Neyshabur, Orhan Firat
Abstract
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bd10286-f2fd-4f7e-8c2a-dbc9412e9029Cited by top-tier papers25
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 271 citations
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 171 citations
- The Unreasonable Effectiveness of Few-shot Learning for Machine TranslationXavier Garcia, Yamini Bansal, Colin Cherry, George F. Foster et al.ICML 2023 · 133 citations
- Getting ViT in Shape: Scaling Laws for Compute-Optimal Model DesignIbrahim M. Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, Lucas BeyerNeurIPS 2023 · 122 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Accuracy on the Line: on the Strong Correlation Between Out-of-Distribution and In-Distribution GeneralizationJohn Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa et al.ICML 2021 · 323 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine TranslationJungo Kasai, Nikolaos Pappas, Hao Peng, James Cross et al.ICLR 2021 · 154 citations
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
Related papers
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- Scaling Laws for Multilingual Neural Machine TranslationPatrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag et al.ICML 2023 · 37 citations
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model PretrainingAnirudh Subramanyam, Yuxin Chen, Robert L. GrossmanICLR 2026 · 6 citations
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas et al.ICLR 2025
- Data Diversification: A Simple Strategy For Neural Machine TranslationXuan-Phi Nguyen, Shafiq R. Joty, Kui Wu, Ai Ti AwNeurIPS 2020 · 75 citations
