Lost in Translation, and Found: Detecting and Interpreting Translation Effects
Shira Wein, Anna Serbina, Jiyuan Ji, Nathan Wolf, Jason DeGraaff, Prajakta Kini, Maria Leonor Pacheco
Abstract
Translationese refers to the statistical patterns that distinguish translated texts from original texts, which are often subtle and imperceptible to human readers. When translated texts appear in either training or testing data, these patterns can negatively affect model performance or warp model evaluation. We approach the task of discerning whether a text was originally written in English or translated into English by fine-tuning contemporary foundation models at distinct item lengths and achieve state-ofthe-art performance (94% Macro F1). Given that these linguistic cues are subtle and often imperceptible to humans, we analyze the features which enable our model's high performance. Employing a suite of interpretabilitybased techniques, we find that: (1) our high accuracy is enabled by a collection of linguistic features, a number of which correspond with linguistic theories of translationese, and (2) pretrained neural models are adept at picking up these features without any fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b28be53-4018-4c5d-8ca7-ab600156c0a5Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
Related papers
- Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese ClassificationDaria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith et al.EMNLP 2021 · 9 citations
- Translating away Translationese without Parallel DataRricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van GenabithEMNLP 2023
- Lost in Literalism: How Supervised Training Shapes Translationese in LLMsYafu Li, Ronghao Zhang, Zhilin Wang, Huajian Zhang et al.ACL 2025 · 12 citations
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang et al.EMNLP 2025
