Translationese as a Language in "Multilingual" NMT
Parker Riley, Isaac Caswell, Markus Freitag, David Grangier
Abstract
Machine translation has an undesirable propensity to produce "translationese" artifacts, which can lead to higher BLEU scores while being liked less by human raters. Motivated by this, we model translationese and original (i.e. natural) text as separate languages in a multilingual model, and pose the question: can we perform zero-shot translation between original source text and original target text? There is no data with original source and original target, so we train sentence-level classifiers to distinguish translationese from original target text, and use this classifier to tag the training data for an NMT model. Using this technique we bias the model to produce more natural outputs at test time, yielding gains in human evaluation scores on both accuracy and fluency. Additionally, we demonstrate that it is possible to bias the model to produce translationese and game the BLEU score, increasing it while decreasing human-rated quality. We analyze these models using metrics to measure the degree of translationese in the output, and present an analysis of the capriciousness of heuristically-based train-data tagging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57ec968a-7942-44cb-a902-4cefdf9391d7Cited by top-tier papers14
- Scaling Laws for Neural Machine TranslationBehrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna et al.ICLR 2022 · 130 citations
- Zero-Shot Cross-lingual Semantic ParsingTom Sherborne, Mirella LapataACL 2022 · 32 citations
- Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPZhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya et al.EMNLP 2021 · 20 citations
- MILCO: Learned Sparse Retrieval Across Languages via a Multilingual ConnectorThong Nguyen, Yibin Lei, Jia-Huei Ju, Eugene Yang et al.ICLR 2026 · 16 citations
- Lost in Literalism: How Supervised Training Shapes Translationese in LLMsYafu Li, Ronghao Zhang, Zhilin Wang, Huajian Zhang et al.ACL 2025 · 12 citations
Builds on2
Related papers
- Translating away Translationese without Parallel DataRricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van GenabithEMNLP 2023
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- Automatic Machine Translation Evaluation in Many Languages via Zero-Shot ParaphrasingBrian Thompson, Matt PostEMNLP 2020 · 7 citations
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang et al.EMNLP 2025
