Scaling Laws for Downstream Task Performance in Machine Translation
Berivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas, Sergei Vassilvitskii, Sanmi Koyejo
Abstract
Scaling laws provide important insights that can guide the design of large language models (LLMs). Existing work has primarily focused on studying scaling laws for pretraining (upstream) loss. However, in transfer learning settings, in which LLMs are pretrained on an unsupervised dataset and then finetuned on a downstream task, we often also care about the downstream performance. In this work, we study the scaling behavior in a transfer learning setting, where LLMs are finetuned for machine translation tasks. Specifically, we investigate how the choice of the pretraining data and its size affect downstream performance (translation quality) as judged by: downstream cross-entropy and translation quality metrics such as BLEU and COMET scores. Our experiments indicate that the size of the finetuning dataset and the distribution alignment between the pretraining and downstream data significantly influence the scaling behavior. With sufficient alignment, both downstream cross-entropy and translation quality scores improve monotonically with more pretraining data. In such cases, we show that it is possible to predict the downstream translation quality metrics with good accuracy using a log-law. However, there are cases where moderate misalignment causes the downstream translation scores to fluctuate or get worse with more pretraining, whereas downstream cross-entropy monotonically improves. By analyzing these, we provide new practical insights for choosing appropriate pretraining data. * Work done when all the authors were at Google. 1 We use the term downstream to refer to the finetuning task or metrics computed on it, and the term upstream to refer to the metrics computed on the pretraining dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Scaling Laws for Optimal Data MixturesMustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier et al.NeurIPS 2025 · 54 citations
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language ModelsChangxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu et al.ICLR 2026 · 45 citations
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang et al.ACL 2026 · 17 citations
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Transfer Learning in Infinite Width Feature Learning NetworksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICLR 2026 · 2 citations
Builds on18
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- LEEP: A New Measure to Evaluate Transferability of Learned RepresentationsCuong V. Nguyen, Tal Hassner, Matthias W. Seeger, Cédric ArchambeauICML 2020 · 279 citations
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 253 citations
Related papers
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge et al.ICML 2025
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li et al.ICML 2025
- Data Scaling Laws in NMT: The Effect of Noise and ArchitectureYamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang et al.ICML 2022 · 61 citations
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of MultilingualityShayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I-Hung Hsu et al.ICLR 2026 · 20 citations
