Statistical Power and Translationese in Machine Translation Evaluation
Yvette Graham, Barry Haddow, Philipp Koehn
Abstract
The term translationese has been used to describe features of translated text, and in this paper, we provide detailed analysis of potential adverse effects of translationese on machine translation evaluation. Our analysis shows differences in conclusions drawn from evaluations that include translationese in test data compared to experiments that tested only with text originally composed in that language. For this reason we recommend that reverse-created test data be omitted from future machine translation test sets. In addition, we provide a reevaluation of a past machine translation evaluation claiming human-parity of MT. One important issue not previously considered is statistical power of significance tests applied to comparison of human and machine translation. Since the very aim of past evaluations was the investigation of ties between human and MT systems, power analysis is of particular importance, to avoid, for example, claims of human parity simply corresponding to Type II error resulting from the application of a low powered test. We provide detailed analysis of tests used in such evaluations to provide an indication of a suitable minimum sample size for future studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ec3abaf-c207-4d90-a424-d48a61592cb8Cited by top-tier papers18
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- Data Scaling Laws in NMT: The Effect of Noise and ArchitectureYamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang et al.ICML 2022 · 61 citations
- Examining Scaling and Transfer of Language Model Architectures for Machine TranslationBiao Zhang, Behrooz Ghorbani, Ankur Bapna, Yong Cheng et al.ICML 2022 · 30 citations
- Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPZhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya et al.EMNLP 2021 · 20 citations
Builds on1
Related papers
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- Revisiting Machine Translation for Cross-lingual ClassificationMikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan et al.EMNLP 2023 · 10 citations
- Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine TranslationXabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy WayACL 2020 · 1 citation
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang et al.EMNLP 2025
